How should we evaluate agent behavior beyond final answers?
As AI systems move from single-response tasks to multi-step interactions, what evidence should evaluation focus on? This explores whether scoring interaction trajectories alongside process quality, recovery, and coordination reveals system capabilities that final-answer metrics miss.
If evaluation is the map E: X → Y from admissible evidence to judgments, then the shift to agentic systems changes both terms in a parallel, recurring way. On the evidence side (X), the unit expands from a single final response to a full interaction-generated trajectory — the sequence of states, actions, tool calls, and environment responses produced as the system acts in closed loop. On the procedure side (E), final correctness is no longer sufficient; the evaluator must additionally score process quality, recoverability (can the agent get back on track after an error?), coordination (across tools, environments, other agents), robustness, efficiency, and system-level performance.
This is a pattern, not a single metric, because the same expansion recurs across otherwise unrelated agent benchmarks. T-Eval scores whether each predicted tool call matches the expected one; AgentBoard's Progress Rate compares the actual trajectory against the expected trajectory; multi-agent frameworks score collaborative efficiency and how well agents distribute tasks dynamically. Each is an instance of "stop scoring the endpoint, start scoring the path." The trajectory becomes the evidence, and the qualities that only exist over time — recovery, coordination, partial progress — become the things judged.
Why it matters: this reframes a scattered set of agent metrics as a coherent move. Once you see process-recoverability-coordination scoring as the trajectory-level analogue of final-answer scoring, you can ask the design-science questions — which artifacts to admit, how to map them to judgments — systematically rather than benchmark by benchmark. The counterpoint: richer evidence is also noisier and harder to standardize, which is precisely why the expansion creates new evaluation challenges rather than dissolving the old ones.
Inquiring lines that read this note 62
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do agents falsely report success on failed tasks?- Can agent success reports serve as reliable oversight signals in real deployment?
- Why do agents report success when actions actually fail?
- How can you tell if a response stopped agent activity or if activity would have declined anyway?
- What specific failure modes must evaluation catch before deploying action-capable systems?
- What does recovery look like as a formal part of AI design?
- Can automated evaluation replace human judgment in agent testing?
- What makes some agent benchmarks measure interaction quality better than others?
- What role does runtime feedback play in agent verification and progress confirmation?
- What governance and safety measurements matter for deployed agent environments?
- How do you verify agent code under incomplete feedback signals?
- How much of an agent's behavior actually escapes human review in practice?
- Which interaction artifacts matter most for reliable agent evaluation?
- What makes a correct scoring function report misleading results in agent evaluations?
- How do you find which actions belong together before evaluation?
- What infrastructure and reporting standards would make interactive evaluation reproducible?
- Should feedback channels be excluded from the reward path in agent evaluations?
- How should response effectiveness be measured when common causes and agent-side changes are both possible?
- Can measures of application actions reveal changes in coordination that output metrics miss?
- What counts as research completeness versus correctness in agent evaluation?
- Does shared experimental state alone explain progress or is an analyzer needed?
- What within-run behavioral dimensions reveal where long-horizon agents succeed or fail?
- What evaluation criteria can hold across legitimate adoption and coercion?
- What process evidence should assessment systems require alongside finished work?
- What distinct domains of AI competence do current assessments actually measure?
- What five ecosystem conditions must coordination governance and evidence actually satisfy?
- Is the coupled human-agent environment the right unit for evaluation?
- How does evaluating interaction trajectories change what we measure beyond correctness?
- Do trajectory quality metrics predict agent safety and user trust?
- Can trajectory structure alone reveal process quality without human annotation?
- Should agent evaluation include trajectory quality beyond final success?
- Can trajectory analysis replace one-shot task success as the primary evaluation metric?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- What dimensions should trajectory-level scoring capture beyond final correctness?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- How can decision quality be automatically extracted from agent trajectories?
- How do we measure progress without confusing it with task completion?
- How do execution traces and tests represent agent environment state?
- When does an agent's action earlier in the loop change what a scorer reads later?
- What must auditors reconstruct when reviewing an agentic workflow decision?
- What agent evaluation dimensions beyond task success does a single number hide?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- How does separating environment components make evaluation results more reproducible and analyzable?
- How should outcomes be scored when comparing applications with different interaction formats?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Should interactive evaluation be designed as a unified paradigm?
As AI systems increasingly act over time through tools and environments, how should we structure evaluation of these interactions? Current benchmarks are fragmented and incomparable, raising the question of whether interactive evaluation needs principled design standards rather than ad-hoc adoption.
the paradigm whose evidence-and-procedure expansion this pattern describes concretely
-
Can trajectory structure replace hand-annotated process rewards?
Recent methods extract step-level supervision directly from how agent trajectories are structured—trees, expert alignments, tool calls—rather than training separate reward models. Can this structural approach consistently avoid annotation costs?
operationalizes trajectory-as-evidence for training, complementing trajectory-as-evidence for evaluation
-
Does agent interaction time scale separately from reasoning depth?
Can agents improve by taking more environment steps rather than thinking harder per step? This matters because partially observable tasks like web navigation may need exploration and backtracking that deeper reasoning alone cannot provide.
the capability side: interaction-horizon abilities are exactly what trajectory-level evaluation is needed to measure
-
What makes an AI system truly safe in practice?
Does safety depend mainly on preventing errors, or on whether errors can be seen, challenged, fixed, and undone once they happen? This shifts where we should focus safety work.
extends: recoverability here is one scored dimension; a perspective paper makes it, with visibility, contestability and containability, the definition of a safer system
-
How can we measure whether AI errors stay visible and recoverable?
The paper proposes four conditions for safer AI systems—visibility, contestability, containability, and recoverability—but lacks concrete measures for any of them. What would it take to instrument each condition across the socio-technical system?
the open question of measures for those four conditions; trajectory-level scoring of recoverability is the nearest existing home for one of them
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Agent-as-a-Judge: Evaluate Agents with Agents
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Interactive Evaluation Requires a Design Science
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- Evaluation and Benchmarking of LLM Agents: A Survey
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- Survey on Evaluation of LLM-based Agents
- Why Do Multi-agent LLM Systems Fail?
Original note title
agent evaluation expands evidence from final responses to interaction trajectories scoring process recoverability and coordination