SYNTHESIS NOTE
Topics›Evaluations›this note

Should interactive evaluation be designed as a unified paradigm?

As AI systems increasingly act over time through tools and environments, how should we structure evaluation of these interactions? Current benchmarks are fragmented and incomparable, raising the question of whether interactive evaluation needs principled design standards rather than ad-hoc adoption.

Synthesis note · 2026-05-28 · sourced from Evaluations

AI evaluation is undergoing a structural change: models are increasingly deployed as systems that act over time through tools, environments, users, and other agents. Yet most evaluation practice still inherits response-centered assumptions — fixed inputs, isolated outputs, a judgment made from a single response. Interactive benchmarks have proliferated, but the landscape is fragmented: they disagree on what interaction artifacts they admit, how trajectories are scored, and what claims their results support. This paper's position is that interactive evaluation should be treated as a principled paradigm, not as the next family of agent benchmarks to collect.

The argument turns on a definition: evaluation is an autonomous mapping E: X → Y from admissible evidence X to judgments Y. Interactive evaluation changes both sides. The evidence X expands from final responses to interaction-generated trajectories; the procedure E must assess not just final correctness but process quality, recoverability, coordination, safety, efficiency, and robustness. From this the authors build a two-axis taxonomy (what artifacts enter; how they map to judgments), derive design principles and reporting standards, and locate where current benchmarks concentrate and what they miss.

Why it matters: the distinction between designing and adopting is the whole point. Adopting interactive benchmarks one at a time produces incomparable, non-reproducible, non-extensible scores — the same fragmentation that plagued early benchmark culture, now at the trajectory level. Treating interactive evaluation as a design science forces explicit protocols, richer trajectory measures, shared infrastructure, and reporting standards that make scores interpretable. The counterpoint the paper concedes: response-centered evaluation remains useful — it is insufficient, not wrong — so the paradigm shift is additive, expanding what counts as evidence rather than discarding the old measures.

Inquiring lines that read this note 14

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does evaluation scope and dimensionality affect what we measure? What prevents conversational agents from taking initiative in dialogue? How does the generation-verification gap limit what we can measure about AI reasoning? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? How do we enforce security boundaries in evaluation environments? What should agent evaluation prioritize to reveal reliable behavior? Should agents decouple planning from perception grounding for better performance? Do reasoning benchmarks predict model performance in long-horizon workflows? Why do persona simulations fail to predict authentic user behavior? When do multi-agent systems outperform single frontier models?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 155 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

interactive evaluation must be designed as a paradigm not adopted as the next benchmark format