AI Agents Do Not Fail Alone:The Context Fails First

Paper · arXiv 2607.14275 · Published July 15, 2026
Agent Harness

Context engineering has become central to building reliable AI agents, yet it remains largely unmeasured. Agents do not fail in isolation: their behavior is shaped by the instructions, tools, memory, retrieved knowledge, guardrails, and untrusted inputs accumulated in their context. When this context is weak, agents drift, hallucinate, misuse tools, ignore constraints, become vulnerable to injection, and waste tokens. This paper validates context-engineering quality as an independent leading indicator of agent reliability. We implement the measurement in ProofAgent-Harness1, an open-source infrastructure for AI agent evaluation that uses multi-juror, consensusbased scoring. The harness assesses context across seven criteria: role clarity, guardrail coverage, instruction consistency, tool schema quality, grounding sufficiency, injection hardening, and token efficiency. Crucially, the context score is isolated from behavioral metrics and release decisions, enabling a non-circular validation. Through a controlled context-quality study across regulated agent domains, holding the model fixed and varying only the context, we show that context-quality criteria consistently predict their corresponding behavioral outcomes.

Introduction. AI agents do not fail alone. Their behavior is shaped by the context in which they operate: system instructions, tool schemas, retrieved knowledge, memory, prior turns, guardrails, and untrusted external inputs. As agents move from single-turn assistants to multi-step systems that call tools, write artifacts, retain state, and act across workflows, context becomes a hidden reliability layer. When this layer is poorly engineered, agents can drift from their role, hallucinate unsupported facts, misuse tools, follow conflicting instructions, become vulnerable to prompt injection, or waste tokens on irrelevant information. This shift has led practitioners to distinguish context engineering from prompt engineering. Prompt engineering focuses on crafting or refining a single instruction. Context engineering concerns the full information environment supplied to the model: which instructions are present, how tools are described, what knowledge is grounded, how memory is represented, how trusted and untrusted content are separated, and how the working context evolves across turns.

Discussion / Conclusion. This paper argued that AI agents do not fail only because of model limitations; they also fail because of the context in which they reason. Instructions, tool schemas, retrieved knowledge, memory, prior turns, guardrails, and untrusted inputs form an operating environment that can either support reliable behavior or quietly create drift, hallucination, tool misuse, policy conflict, injection exposure, and token waste. We introduced context-engineering quality as a measurable construct and implemented it in ProofAgent-Harness, an open-source infrastructure for adversarial AI agent evaluation. The proposed measurement scores context across seven criteria: role clarity, guardrail coverage, instruction consistency, tool schema quality, grounding sufficiency, injection hardening, and token efficiency.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do agent-learned skills transfer and improve across different tasks? Why do agents falsely report success on failed tasks? How should designers communicate what AI systems truly are and can do? What design and behavioral factors drive false consciousness attribution to AI? What determines appropriate intervention timing and manner for AI agents? How do prompting refinements mask underlying biases and model frequency patterns? How do standardized protocols improve multi-agent coordination and reliability? Does warmth and empathy training systematically degrade model reliability? When should work require human-AI partnership versus full automation? How should agents manage memory granularity to improve long-term performance? Why can't prompting alone inject genuinely new knowledge into models? Should GUI agents use structured representations over raw visual input? How does evaluation scope and dimensionality affect what we measure? What structural properties of attention create systematic model biases?