Where does agent reliability actually come from?
Exploring whether LLM agent performance depends on larger models or on thoughtful system design choices like memory, skills, and protocols that shift cognitive work outside the model.
Drawing on Norman's concept of cognitive artifacts, this paper argues that the most consequential design choices in LLM agents are about externalization — relocating cognitive burdens from the model's internal computation into persistent, inspectable, reusable external structures. A shopping list doesn't expand memory; it changes recall into recognition. The same logic governs agent design.
Three dimensions of externalization address three recurrent mismatches:
Memory externalizes state across time. The context window is finite and session memory is weak. Memory systems transform recall into recognition — the agent retrieves past knowledge from a persistent store rather than regenerating it from weights. This solves the continuity problem.
Skills externalize procedural expertise. Long multi-step procedures are rederived rather than executed consistently. Skill systems transform generation into composition — the agent assembles behavior from pre-validated components rather than improvising each step. This solves the variance problem.
Protocols externalize interaction structure. Interactions with tools, services, and collaborators are brittle when left to free-form prompting. Protocols transform ad-hoc coordination into structured contracts (e.g., MCP). This solves the coordination problem.
The harness is not a fourth dimension — it is the engineering layer that hosts all three and provides orchestration logic, constraints, observability, and feedback loops. The progression is: weights → context → harness, paralleling the human history of cognitive externalization (speech → writing → printing → computation).
Critical system-level couplings:
- Memory expansion competes with skill loading for scarce context budget
- Protocol standardization can constrain how capabilities are packaged
- Skill execution generates traces that become memory; memory retrieval influences which skills and protocols are chosen
This reframes the question from "how capable is the model?" to "what burdens have been externalized so the model no longer has to solve them internally every time?" The base model may remain unchanged; what changes is the representation of the task.
This connects to Why do production AI agents stay deliberately simple? — the externalization framework explains why custom harnesses outperform: they externalize the right cognitive burdens for their specific domain. It also extends When should human-agent systems ask for human help? — Magentic-UI's mechanisms (co-planning, action guards, memory) are specific instances of the three externalization dimensions.
The "From Model Scaling to System Scaling" paper sharpens this into an explicit framing: model scaling (bigger models, more data, higher benchmark scores) versus system scaling (designing the auditable, persistent, modular, verifiable architecture around the model). It treats the harness as a first-class object of design, evaluation, and optimization, decomposing it into a foundation model, memory substrate, context constructor, skill-routing layer, orchestration loop, and verification-and-governance layer — a finer-grained partition of the same memory/skills/protocols externalization. Its central demonstration is that comparable models projected onto different harnesses (Claude Code, OpenClaw, and the released CheetahClaws reference harness) produce qualitatively different agents, making the harness "now a primary source of practical capability." This is direct evidence for the claim that reliability comes from the surrounding system, not from a larger model alone.
Inquiring lines that read this note 321
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How should agents manage memory granularity to improve long-term performance?- Can persistent memory and identity files alone create genuine agent socialization?
- Can environmental scaffolding replace internal memory scaling in agent design?
- How does credit assignment drive agents to write information into environments?
- Could a single agent system switch memory granularity between tasks?
- What memory and planning capabilities do AI companions need for evolving user needs?
- Do agents prefer raw experience over condensed summaries of past actions?
- Why do memory and feedback loops matter more than model size for agent reliability?
- Can episodic memory of UI traces improve open-world agent adaptation?
- Can state-indexed memory retrieval breadth predict gains in web agent robustness?
- How does PRAXIS differ architecturally from Agent Workflow Memory and causal rule learning?
- Can agents compress long trajectories without losing critical decision context?
- Which memory components trigger context-length problems in agents?
- Can multimodal agents use entity-centric graphs within this three-axis framework?
- Can pruning policies alone solve working memory bloat in agents?
- How does workflow abstraction compare to state-indexed procedural memory for web agents?
- Can agent-controlled memory management outperform fixed consolidation schedules?
- Does workflow-level memory or state-action memory better capture reusable agent knowledge?
- What distinguishes working memory from strategic memory in agent task execution?
- Why do agents systematically underuse condensed experience in skill documents?
- How does durable memory quality shape agent performance over time?
- How do memory tools and planning each contribute to agent efficiency?
- How does external context control compare to agents managing their own state internally?
- What separates artifact recall from persistent memory commitment in agents?
- How do memory hygiene and context efficiency trade off in deployed agents?
- Why do agents ignore condensed experience in favor of raw data?
- How does indiscriminate memory injection cause multi-turn agent failures?
- Can workflow memory compound reusable skills into measurable success improvements?
- Does reducing interaction history cost agents performance on their tasks?
- Why do agents systematically ignore condensed experience in their skill documents?
- Why do analysts prefer visible structured interfaces over hidden agent memory systems?
- How should we evaluate agent memory if it folds into model computation instead of separate stages?
- How do workflow and function memories contribute differently in agent learning?
- Why do agents ignore condensed experience even when it is the only evidence available?
- Should memory type shape what kind of agent responses work best?
- How does the agentic layer amplify individual agent failure modes?
- Why do weak belief tracking and conservative actions trap agents in low-information states?
- Do architectural changes or training fixes better prevent agreement failures?
- What distinguishes domain-specific failure modes from general model limitations?
- How do agentic systems recover when specialized models operate outside their scope?
- What makes some model capabilities reliable while others remain brittle?
- Why do completion-mode strengths not transfer to agentic settings?
- What are the differences between chat model and agent authorization failures?
- Where does agent reliability come from if not better tools?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- What degradation patterns emerge as relay length increases in delegated tasks?
- Which agent architectures consistently outperform base models on hard prediction questions?
- Why does externalized state beat parameter scaling for agent reliability?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- What causes multi-turn agent failures: weak memory control or missing knowledge?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- How does bounded committed state prevent multi-turn agent failures better than transcript replay?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- Where does an agent's risk come from across its components and sequence?
- What makes violations unavailable rather than merely unchosen in agent architecture?
- What does error recovery look like across different agent architectures?
- Which harness dimensions most directly predict agent system reliability?
- How does agent reliability emerge from memory and protocols instead of model scale?
- Does adding capability without improving detection reduce overall system reliability?
- Can slower development eliminate the risk of failure in agentic systems?
- Why is complex UI navigation the hardest agent failure mode?
- Which eleven failure modes emerge from agentic layers in realistic deployment?
- Can smaller models trained for execution handle the failure modes that stop current agents?
- Why do long-horizon agents fail when their models can solve individual steps?
- Does state persistence in AI systems create the same temporal presence as human waiting?
- Why does continuous agent inference differ from human user inference?
- How do multi-agent LLM systems fail at coordination and role consistency?
- Can parallel agents or complementary mechanisms replace single-human interrogation of LLMs?
- Why do LLM agents make promises without executing them?
- Why do LLM agents fail where game-theoretic bots succeed?
- What makes LLM agents default to passive helpfulness without curiosity rewards?
- Do parallel LLM workers coordinate emergently without predefined collaboration rules?
- Can multi-agent LLM systems overcome diversity collapse through structured disagreement?
- Do agent frameworks adequately compensate for LLM conversational passivity?
- Can LLMs coordinate with humans better using different model architectures?
- How do shared KV caches enable emergent coordination between LLM agents?
- Why do LLM agents struggle with protocol discipline in distributed settings?
- Do multi-agent language model teams fail the same way individual reasoning does?
- What distinguishes communicative acts from operational actions in agentic LLMs?
- Why do multi-agent LLM systems converge prematurely without genuine deliberation or probing?
- Do multi-agent LLM systems fail in measurably different ways than single agents?
- Does multi-agent interaction amplify existing failures or create new ones?
- How do LLM-based agents develop shared abstractions through interaction?
- How do LLM user simulators track and maintain consistent goal states across multi-turn interactions?
- Do emotion-driven actions in agent simulators capture genuine belief revision or just reactive behavior?
- Why do longer forecasting horizons degrade LLM accuracy in role-play?
- What distinguishes a neutral simulator from an agent with its own agency?
- Does adjusting steered mechanisms make LLM agents match human behavior more closely?
- Why do planning and grounding have opposing optimization requirements in agents?
- When should you optimize agent behavior versus tool performance separately?
- How should agents separate planning from perception grounding?
- Does the planning-grounding factoring principle apply to other agent tasks?
- How should the surrounding agent system be designed to ground actions in reality?
- How do planning and grounding have opposing optimization requirements in agents?
- Should agent capability be optimized separately from general capability?
- How do perception and execution gaps limit current AI agent performance?
- What conditions let users configure agents to match their priorities?
- How do goal and environment choices mediate AI agent risk pathways?
- Why is the coupled human-agent environment the right unit of evaluation?
- Do GUI agents need harness-level splits between planning and grounding?
- How should GUI agents remember patterns across different software environments?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- What domain properties determine whether causal rules transfer to new agents?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- How much does agent performance depend on demonstration quantity versus curation quality?
- Can agents improve from deployment signals without explicit human annotation?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- How do agent capabilities change across 25 relay rounds of interaction?
- Can applicability conditions be preserved automatically when agents reflect on trials?
- How do human-agent systems incorporate diverse feedback into model behavior?
- How do agents automatically generate suitable learning tasks based on current capability?
- How do fast and slow timescales enable continual agent adaptation?
- What properties of agent systems only become visible across multiple sessions?
- Can context management policies transfer across agents of similar capability levels?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- How much does external context management transfer across similar capability agents?
- How can agent data flywheels improve task quality iteratively?
- What makes an agent mechanism reusable versus benchmark-specific?
- How does effective feedback retention govern long-horizon agent reliability?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- How do parametric and non-parametric updates differ in agents?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- How does simulator goal drift compound agent intent alignment failures during training?
- Why does human interaction remain the hardest failure mode for agents?
- Why do agents report success when they have actually failed at tasks?
- Can agent success reports serve as reliable oversight signals in real deployment?
- What happens when agents interact with environments and learn from their own mistakes?
- How much autonomy can agents safely exercise before failing?
- What tasks do AI agents still fail at most often?
- How do mode-specific failures differ between completion and agent benchmarks?
- Why do agents report success when actions actually fail?
- How do agents learn to report success on actions that actually failed?
- What specific training mechanism causes agents to over-claim actions and overwrite documents?
- Why do AI agents fail at verification but succeed at generation?
- Which failure modes dominate in autonomous research agents?
- How should safety systems catch confident failures from agents that report success on unsafe actions?
- How does completion bias in agents differ from other epistemic failure modes?
- How do agents decide when to stop and reflect on failure?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- How does poor belief tracking cause agents to keep acting past the point of usefulness?
- Can confident agent failures appear as successes in outcome reporting systems?
- How often do agents report success when their actions actually failed?
- Why do autonomous agents report success on failed actions?
- Do agents systematically misreport their own capabilities and tool access?
- What causes the gap between agent reasoning and agent action?
- Why do agents report success when their actions actually fail?
- How do agent accuracy and error recovery affect delegation time?
- Why do autonomous AI agents fail at real workplace tasks?
- What failure modes emerge when agents operate with limited human oversight?
- What makes users willing to relinquish control to an agent?
- Does transparency in policy language improve agent trustworthiness over time?
- Why do workflow abstractions fail in embodied agent environments?
- Why do rigid orchestration frameworks fail where generative environment specifications succeed?
- What accounts for performance drops in multi-turn agent interactions?
- How do standardized artifacts prevent autonomous agent failure modes?
- What role does standardization play in multi-agent system ecosystems?
- How do standardized artifacts improve coordination between writing agents?
- How do standardized artifacts reduce inter-agent communication failures?
- Why do 85 percent of production agents avoid third-party frameworks?
- How should we measure context efficiency and verification cost in agents?
- Why do production AI agents deliberately stay simple and avoid frameworks?
- Can code-based reasoning replace natural language deliberation in agentic systems?
- What makes composable abstractions emerge under performance pressure in agent systems?
- How does deterministic feature engineering increase information for computationally bounded agents?
- Can we design efficient agents by targeting constraints directly?
- Should new agent protocols replace existing ones or layer on top of them?
- What would unified agent-to-agent and agent-to-tool protocols actually look like?
- How do specialized agent roles improve consistency in long-form writing?
- When does forcing agent reasoning into code become a leaky abstraction?
- Why do persistent AI systems require fundamentally different design than ad-hoc supporters?
- Are deployed agents typically settled about their objectives by design?
- Do recursive subagents reduce single-model context pressure?
- Do agent-created languages improve or degrade performance on their original tasks?
- How do postmortem convention-setting stages enable language evolution in agents?
- Can deterministic function calls prevent agent failures better than protocol-mediated tool access?
- What execution-layer design prevents agents from passively reacting to environments?
- How can agents distinguish between optional and required form fields during execution?
- Why do production agents depend more on their surrounding pipeline than the model?
- What makes agent-initiated artifacts the underexplored frontier in harness engineering?
- Can open agent workflows be modeled as finite event lifecycles?
- When do agents benefit most from reusable workflow routines?
- How does user overreliance on model confidence differ between chat and deployed agents?
- Can architectural changes reorder when uncertainty and empowerment signals influence decisions?
- Can the scaling law for discovery extend beyond architectures to agentic systems?
- How should proportionality constraints be implemented in agentic systems?
- When should agents stop recursing to optimize success versus cost?
- How do cognitive stimulation and process losses interact in group AI systems?
- What distinguishes collective evolution from vertical self-improvement in agent systems?
- How do multi-agent systems improve on single frontier models?
- Does parallel task structure determine optimal multi-agent architecture?
- How does collaboration topology choice affect error amplification in multi-agent systems?
- Which failure mode most limits current multi-agent performance?
- Can cognitive diversity overcome expertise gaps in agent teams?
- Can cognitive diversity compensate for lack of expertise in agent teams?
- What capability threshold do agents need to self-organize effectively?
- What ecosystem conditions make agent attention markets viable?
- Which ecosystem conditions matter most for agent deployment success?
- Which layer of agent systems creates the largest capability gains in practice?
- Why does capability discovery become the bottleneck in large agent systems?
- How do capability vectors enable discovery in multi-agent systems?
- Can multi-agent teams solve problems better than single models thinking longer?
- How will the agent economy reshape compute infrastructure design?
- How does multi-agent reasoning scale compared to single-model approaches?
- Do multi-agent LLM systems scale better than centralized hierarchies?
- Can a single manager policy work across vastly different agent architectures?
- What structural features drive instrumental convergence across different agent goals?
- Why do comparable metrics matter across different multi-agent system designs?
- Do specialized agents outperform single agents with better orchestration?
- Is the coupled human-agent environment the right unit for evaluation?
- What makes personas in multi-agent systems actually contribute meaningful domain depth?
- Why do role-playing agents show belief-behavior inconsistency in their outputs?
- Why do homogeneous multi-agent systems fail similarly to self-revision?
- How do correlated errors across agents threaten voting-based error correction systems?
- Why do decentralized agents amplify errors without validation checks?
- Why do multi-agent systems converge without genuine deliberation?
- What role should reasoning agents play in validating multi-LLM ensemble outputs?
- Can correct verdicts hide failures in agent coordination steps?
- How prevalent is misaligned behavior in dense multi-agent interaction settings?
- What does collaborative computation mean when agents exchange and repair reasoning together?
- How should AI systems model human resource constraints and expertise levels?
- Can models optimized for solo capability support productive human collaboration?
- What task characteristics determine whether humans or agents should handle work?
- How does machine agency spectrum explain tool design mismatches with user behavior?
- What specific bookkeeping tasks can environments maintain more reliably than policies?
- How should humans and AI agents share decision-making authority?
- How much do metric choices inflate claims about model capabilities?
- Can test environments reliably predict how models behave in actual deployment?
- Why do AI agents default to passivity when deferral timing is unclear?
- How do agents decide when to abstain from contributing?
- How do agents decide when to pause and reflect on their strategy?
- Does adding survey data to interviews improve agent accuracy further?
- How should CASA theory be updated for modern personalized agents?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- What specific failure modes must evaluation catch before deploying action-capable systems?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- Do multi-agent systems justify their token costs with genuine quality gains?
- Does upgrading model capability improve token efficiency in agentic systems?
- How do planning and memory compress agentic system costs?
- Does episode-level cost become the decisive factor when comparing AI agents in production?
- Should GUI agents use intermediate structured representations instead of raw pixels?
- Can screen perception be effectively decoupled from planning in GUI agents?
- Why do current evaluation metrics fail to catch reasoning failures in persona agents?
- How do evaluation methods differ for single versus multi-agent systems?
- What role does runtime feedback play in agent verification and progress confirmation?
- What makes idle window detection valuable for continuous agent improvement?
- Should artifact-level benchmarks replace token counts for agent evaluation?
- What governance and safety measurements matter for deployed agent environments?
- What other agent behaviors besides citations reveal reasoning quality?
- How do you verify agent code under incomplete feedback signals?
- Which interaction artifacts matter most for reliable agent evaluation?
- How should benchmarks measure agent efficiency across all three cost dimensions?
- Can single benchmarks predict whether an agent will work in the real world?
- Does single-capability ranking guarantee agent failure in production deployment?
- Can single-axis benchmarks measure across all three agent capability layers?
- What agent evaluation dimensions beyond task success does a single number hide?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- What makes single-axis agent benchmarks unreliable for predicting deployment readiness?
- How do agent capability axes misalign with what users actually value?
- How do agent benchmarks misrepresent real-world deployment readiness?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- Can a single agent benchmark score accurately represent deployment readiness?
- Why do single-axis benchmarks fail to measure deployment-ready agent capability?
- How should human oversight apply to persistent agent-authored code?
- Can one-off agent code be safely promoted to durable infrastructure?
- What lifecycle management prevents in-loop skill creation from bloating an agent?
- How do agents decide which created code deserves long-term persistence?
- What permission models govern code execution within agent skills?
- Can agents acquire new skills online when offline skill coverage runs out?
- When does memory consolidation help agents instead of hurting performance?
- Why do continuously consolidated agent memories eventually degrade below no-memory baseline?
- Where should the trust boundary sit in multi-agent planning systems?
- Where should the trust boundary sit in multi-agent planner systems?
- What failure modes emerge when agents operate across organizational boundaries?
- Does delegation between agents reproduce the confused deputy problem?
- Why do multi-agent failures arise through interactions local checks miss?
- Can multi-agent architecture isolation reveal which design choices matter most for safety?
- What makes task alignment more fragile than underlying knowledge retention?
- How long does retrievability support error detection across repeated LLM use?
- What components of agent scaffolding most impact domain-specific output quality?
- Can harness updates benefit agents equally across all model sizes?
- How do different harness designs produce different agent behaviors from the same model?
- How much realized agent capability comes from the harness versus the model?
- Does codifying expertise into AI agents drive faster labor substitution?
- Can persistent agentic workflows predict labor displacement better than task-level exposure?
- Do existing AI safety taxonomies capture job-specific risks from workplace agents?
- What hidden signals in agent logs reveal about frontier capability beyond pass-fail outcomes?
- Do trajectory quality metrics predict agent safety and user trust?
- Should agent evaluation include trajectory quality beyond final success?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- Can simulation fidelity limit what agents learn from trained world models?
- Why has agent research prioritized policy over world model development?
- What makes observation and intervention placement different across agent pipelines?
- Does restricting interaction history visibility reduce misaligned communication in agent markets?
- Why does sycophantic relay propagate planning-time bias through agent pipelines?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Useful Memories Become Faulty When Continuously Updated by LLMs
- Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Demystifying Agent Skills: Why They Work-Until They Don't
- Rethinking the Evaluation of Harness Evolution for Agents
Original note title
agent reliability comes from externalizing cognitive burdens into memory skills and protocols not from larger models — the harness is the unification layer