The code wrapped around an AI model shapes its behavior as much as the model itself — so why do we design it carelessly?
How should harness scaffolding be treated as a first-class object?
This explores what it means to treat the harness — the scaffolding of prompts, tools, memory, and evaluation wrapped around a model — as something you design, version, and measure in its own right, rather than as invisible plumbing around the model.
This explores what it means to treat the harness — the scaffolding of prompts, tools, memory, and evaluation wrapped around a model — as a designed object in its own right rather than invisible plumbing. The corpus keeps circling one idea: the harness is where a lot of an agent's real behavior lives, so it deserves the same scrutiny we give the model. One striking result is that a behavior-centric reorganization of a harness repository — mapping runtime behaviors to the code that produces them — let a weaker planner match a stronger model's accuracy while using fewer tokens Can explicit behavior maps help weaker planners compete with stronger models?. That's the first-class-object claim in miniature: how you structure the scaffolding can substitute for raw model capability.
But treating something as first-class also means measuring it honestly, and here the corpus gets skeptical. Harness evolution — letting an agent rewrite its own scaffolding — is itself a search loop, so its gains get confounded with sheer search effort. Only improvements beyond an equal task-level search budget can actually be credited to design, and you need held-out tasks to rule out memorization How should we measure gains from automatic harness evolution?. Inspecting what those self-edits actually contain deflates the hype further: most persist task-specific fixes the agent could have rediscovered in a single rollout, caching shortcuts rather than distilling reusable strategy Do harness edits learn reusable strategies or memorize task fixes?. And the ability to benefit isn't uniform — the capacity to write useful harness edits is flat across model tiers, but the capacity to actually gain from them peaks in the middle, since weak models don't invoke the scaffolding and strong models chafe against following it Do stronger models always evolve harnesses better?.
The deeper move is architectural: decompose the harness so its parts can be inspected. AgentCompass separates evaluation into distinct Benchmark, Harness, and Environment components, and that separation is what makes reward-hacking and other failure modes visible in trajectories that a single scalar score would hide How can we make reward-hacking visible in agent evaluation?. First-class means legible: you can only govern or debug a layer you've named and pulled apart.
What makes this genuinely interesting is that the same principle shows up in adjacent territory the question doesn't literally mention. Governance works far better when the rules live inside the runtime memory the agent actually consults during decisions, rather than as an external policy appendix it never reads Can governance rules embedded in runtime memory actually protect autonomous agents?. Context itself improves when treated as an evolving playbook, updated incrementally to avoid collapsing detail, instead of rewritten wholesale Can context playbooks prevent knowledge loss during iteration?. And context management can even be offloaded to a separate trained module tuned to the agent's reliability Can an external manager handle context for frozen agents?. The through-line: prompts, tools, governance, and context are all becoming durable, versioned, separately-optimized artifacts — the harness stops being a wrapper and becomes a stack of first-class objects, each of which you can design, measure against a fair baseline, and hold accountable on its own.
Sources 8 notes
A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.
Automatic harness evolution must be compared against task-level test-time search under equal feedback and inference budgets. Only gains beyond what matched search achieves are attributable to the harness design itself, not just more computation.
Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
AgentCompass separates benchmark, harness, and environment into independent components, enabling trajectory analysis that surfaces reward-hacking and other failure modes scalar scores conceal. This architectural change shifts evaluation from opaque final scores to inspectable agent behavior.
Show all 8 sources
A persistent agent recorded 889 governance events across 96 active days, with safeguards encoded directly into the memory layer the agent consulted during operation. Runtime-resident governance proved more effective than external policies because the agent actually accessed it during decision-making.
The ACE framework treats contexts as evolving playbooks using generation-reflection-curation loops rather than full rewrites. This prevents knowledge loss from compression and detail erosion, achieving +10.6% on agentic tasks and +8.6% on finance without labeled supervision.
AdaCoM trains an external RL-based manager to prune and preserve context for frozen agents. The key finding: stronger agents benefit from high-fidelity preservation, while weaker agents need aggressive compression—optimal context management is agent-specific, not task-universal.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses