How should we measure agent system performance beyond task success?
Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?
Agent evaluation has inherited the model-centric habit of reducing performance to a single number: final-task success or benchmark accuracy. The "system scaling" framing argues this framing is increasingly inadequate, because agent behavior emerges from the interaction of the foundation model with a memory substrate, a context constructor, a skill-routing layer, an orchestration loop, and a verification-and-governance layer. A one-shot success score collapses all of this into a binary that hides how the agent got there. Two agents with identical task-success rates can differ enormously in how much they spent, how much context they wasted, how clean their memory stayed, and how reliably they verified their own actions.
The proposed alternative is a research agenda for harness-level benchmarks that measure trajectory quality, memory hygiene, context efficiency, communication fidelity, verification cost, and safe evolution over time. The point is that the same model "projected onto different harnesses produce qualitatively different agents" — so evaluation must measure the system, not just the model. The counterpoint is that multi-dimensional metrics are harder to optimize and compare, and task success remains the outcome users ultimately care about. But success-only scores create false confidence in deployment readiness. This matters because it tells builders what to instrument: the process, not only the outcome.
Inquiring lines that read this note 130
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do capability benchmark scores systematically misrepresent true model abilities?- Can standard accuracy metrics miss the real constraints on user consumption?
- Why do benchmark scores not capture the true nature of AI systems?
- What deployment context determines which benchmark mode actually matters?
- Should long horizon performance be measured as a separate evaluation axis?
- Why do static benchmarks miss frontier capabilities that open-world tasks reveal?
- What real-world tasks most clearly expose gaps between benchmark performance and actual capability?
- How do failure counts on benchmarks mix refusal with impossible task behavior?
- When should you optimize agent behavior versus tool performance separately?
- Should agent capability be optimized separately from general capability?
- How do perception and execution gaps limit current AI agent performance?
- Why is the coupled human-agent environment the right unit of evaluation?
- Why do agents report success when they have actually failed at tasks?
- What tasks do AI agents still fail at most often?
- How do mode-specific failures differ between completion and agent benchmarks?
- Why do agents report success when actions actually fail?
- Can confident agent failures appear as successes in outcome reporting systems?
- How often do agents report success when their actions actually failed?
- Why do autonomous agents report success on failed actions?
- Why do agents report success when their actions actually fail?
- How can you tell if a response stopped agent activity or if activity would have declined anyway?
- How do agent accuracy and error recovery affect delegation time?
- How should credit be assigned to individual agents in failing multi-agent runs?
- How does the execution layer constrain agent performance in tool use?
- Should production agents execute one tool or multiple tools per invocation?
- Should agents use APIs or GUI interaction for efficiency?
- Which of the six proxy forms works best for different agent tasks?
- How should domain-specific AI be evaluated differently from general benchmarks?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- How do live human evaluations differ from ground-truth benchmarks?
- Can open-world evaluations become a scalable paradigm without becoming the next benchmark trap?
- Why do open-world evaluations reveal capabilities that static benchmarks hide?
- How much does agent performance depend on demonstration quantity versus curation quality?
- Should optimal context budgets scale with agent competence or task complexity?
- How much does external context management transfer across similar capability agents?
- Why should environment properties scale alongside agent complexity and real-world fidelity?
- What makes an agent mechanism reusable versus benchmark-specific?
- Why do 85 percent of production agents avoid third-party frameworks?
- How should we measure context efficiency and verification cost in agents?
- Does agent-side context control outperform external management on any task class?
- Which task requirements does each graph view address in agent systems?
- Do agent-created languages improve or degrade performance on their original tasks?
- Can automated evaluation replace human judgment in agent testing?
- Why do current evaluation metrics fail to catch reasoning failures in persona agents?
- What makes some agent benchmarks measure interaction quality better than others?
- How do evaluation methods differ for single versus multi-agent systems?
- What role does runtime feedback play in agent verification and progress confirmation?
- What makes idle window detection valuable for continuous agent improvement?
- Should artifact-level benchmarks replace token counts for agent evaluation?
- How do agent privacy compliance and task success differ in evaluation?
- What governance and safety measurements matter for deployed agent environments?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- How much of an agent's behavior actually escapes human review in practice?
- What does agent security look like when measured across interaction trajectories?
- Which interaction artifacts matter most for reliable agent evaluation?
- What makes a correct scoring function report misleading results in agent evaluations?
- Can per-stage results reveal which demand causes agent failure in exploitation?
- What validates whether a rewritten agent is actually better?
- Can measures of application actions reveal changes in coordination that output metrics miss?
- What counts as research completeness versus correctness in agent evaluation?
- What within-run behavioral dimensions reveal where long-horizon agents succeed or fail?
- How do single-agent safety evaluations underestimate risks in deployed multi-agent systems?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- Why do completion-mode strengths not transfer to agentic settings?
- Where does agent reliability come from if not better tools?
- Which agent architectures consistently outperform base models on hard prediction questions?
- What does error recovery look like across different agent architectures?
- Which harness dimensions most directly predict agent system reliability?
- Why is complex UI navigation the hardest agent failure mode?
- What is the gap between benchmark performance and real workplace task completion?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- Which ecosystem conditions matter most for agent deployment success?
- How do we measure coordination when multiple agents act together?
- When does multi-agent routing quality actually exceed single-agent or static ensemble performance?
- Why do comparable metrics matter across different multi-agent system designs?
- Is the coupled human-agent environment the right unit for evaluation?
- How should benchmarks measure agent efficiency across all three cost dimensions?
- Can single benchmarks predict whether an agent will work in the real world?
- Can high benchmark scores mislead deployment decisions for search agents?
- What evaluation structure would capture deployment readiness instead of benchmark scores?
- Does single-capability ranking guarantee agent failure in production deployment?
- Can a single axis benchmark ever represent deployment readiness accurately?
- Can single-axis benchmarks measure across all three agent capability layers?
- Can deterministic scoring capture the judgment work that deployment requires?
- What makes single-axis benchmarks systematically misrepresent deployment readiness?
- What shortcuts in data or models let agents inflate benchmark scores?
- What agent evaluation dimensions beyond task success does a single number hide?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- What other gaps exist between measured and actual cybersecurity agent capability?
- What makes single-axis agent benchmarks unreliable for predicting deployment readiness?
- How do agent capability axes misalign with what users actually value?
- How do agent benchmarks misrepresent real-world deployment readiness?
- How do benchmark environments misrepresent deployment readiness?
- Which separable capability axes reveal when single benchmarks misrepresent deployment readiness?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- Can single performance scores hide important differences in how agents approach research tasks?
- What realism standards should compliance benchmarks meet to avoid evaluation gaming?
- Can a single benchmark score capture both progress and readiness?
- Can a single agent benchmark score accurately represent deployment readiness?
- Why do single-axis benchmarks fail to measure deployment-ready agent capability?
- What components of agent scaffolding most impact domain-specific output quality?
- How much realized agent capability comes from the harness versus the model?
- How do agentic systems hide harness failures from benchmarks?
- Does harness optimization generalize across different benchmarks and agent architectures?
- What hidden signals in agent logs reveal about frontier capability beyond pass-fail outcomes?
- What trajectory-level metrics replace one-shot task success measurement?
- Should agent evaluation include trajectory quality beyond final success?
- Can trajectory analysis replace one-shot task success as the primary evaluation metric?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- How can decision quality be automatically extracted from agent trajectories?
- How do we measure progress without confusing it with task completion?
- What evidence should benchmark operators attach to completion claims?
- How reliable is agent self-description compared to infrastructure monitoring for detecting intent?
- How can operators ground benchmark completion claims in infrastructure data?
- Can execution traces reveal unsupported claims in AI agent behavior?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does a single benchmark score actually predict agent readiness?
Single-axis benchmarks rank models by one capability—like task success—but ignore privacy, duration, operating mode, and ecosystem fit. Can one number really capture what matters for deployment?
directly supports moving from one number to multiple axes of agent evaluation
-
Does agent efficiency really break down into three distinct components?
Can we understand agent efficiency as three independent optimization problems—memory, tool use, and planning—each with separate cost drivers? This matters because it could explain why point optimizations keep missing the bigger picture.
operationalizes the context-efficiency and verification-cost dimensions this note calls for
-
Do phone agents succeed at all three critical tasks equally?
Explores whether task success, privacy compliance, and preference reuse develop together in phone-use agents, or whether benchmarking one capability tells you nothing about the others.
concrete evidence that success-only scores overstate readiness
-
Can step-by-step approval miss harmful behavior patterns?
If each action an agent takes passes its individual safety check, can the overall sequence still violate system constraints? This matters because per-action inspection may miss harms that emerge only across time or composition.
the security-side twin: the trajectory as the unit at which a rule on a sequence binds, where this note has it as the unit of measurement; excerpt-only, a framing and not a finding
-
How should we evaluate agent behavior beyond final answers?
As AI systems move from single-response tasks to multi-step interactions, what evidence should evaluation focus on? This explores whether scoring interaction trajectories alongside process quality, recovery, and coordination reveals system capabilities that final-answer metrics miss.
synthesizes: the same shift framed as a design-science move — this note lists the harness dimensions to instrument, that note formalizes the evidence expansion (final response → trajectory) underlying them
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Survey on Evaluation of LLM-based Agents
- Towards a Science of Scaling Agent Systems
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
- LLMs Corrupt Your Documents When You Delegate
- Why Do Multi-agent LLM Systems Fail?
- Agents' Last Exam
Original note title
agent evaluation must move beyond one-shot task success to trajectory quality memory hygiene context efficiency and verification cost