SYNTHESIS NOTE
Topics›Evaluations›this note

What predicts success in ultra-long-horizon agent tasks?

Does an agent's initial solution quality matter more than its willingness to iterate? AUTOLAB's frontier-model benchmark suggests persistence through feedback loops may be the true differentiator.

Synthesis note · 2026-06-27 · sourced from Evaluations
How does test-time scaling work for individual research agents?

AUTOLAB reframes what a long-horizon agent benchmark should test. Most agentic evals score either single-turn responses or short interactive trajectories; AUTOLAB instead hands the agent a correct but deliberately suboptimal baseline across 36 expert-curated tasks (system optimization, CUDA kernels, model development, puzzles) and asks it to improve the artifact within a strict wall-clock budget. The striking empirical result, across 17 frontier models, is that the dominant predictor of success is not the quality of the agent's initial attempt but its persistence — its willingness to repeatedly benchmark, edit, and incorporate noisy empirical feedback over many cycles. Most models, including proprietary ones, either terminate prematurely or exhaust their budget with minimal progress; claude-opus-4.6 is called out as a strong exception.

This is a sharper, more operational claim than "agents should iterate." It says the binding constraint is a behavioral disposition toward sustained empirical grounding, and that disposition is unevenly distributed across models that look comparable on one-shot benchmarks. It grounds How should we measure agent system performance beyond task success? with a concrete trajectory-level predictor, and it sits naturally alongside Does raw token spending actually predict agent performance? — persistence only pays if each loop returns informative, retained feedback, otherwise it is budget-burning churn, not progress.

The mechanism cuts against itself, however. Do models fail worse when their own errors fill the context? implies that more loops mean more accumulated mistakes in context, which should degrade the very iteration AUTOLAB rewards. The reconciliation is probably that persistence pays only when paired with calibrated scoring that lets the agent see whether an edit actually helped — pure persistence without trustworthy feedback would amplify error. That is why the authors single out harness design as the promising lever: the harness, not the backbone alone, decides whether long horizons compound feedback or compound noise.

Inquiring lines that read this note 78

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What trajectory-level metrics beyond task success best evaluate agent performance? Do language models develop actual world models or merely task heuristics? How does harness optimization generalize across different model architectures and domains? Why do standard benchmarks fail to predict agent deployment success? How do agent-learned skills transfer and improve across different tasks? How should agent systems validate and persist generated code artifacts? How do capability benchmark scores systematically misrepresent true model abilities? How does the generation-verification gap limit what we can measure about AI reasoning? How do standardized protocols improve multi-agent coordination and reliability? Why do agents falsely report success on failed tasks? How can evolutionary algorithms maintain diversity during solution search? How do evaluation practices shape which failures stay visible? How do neighboring agents influence whether others cooperate or collude? Can harness architecture and protocols provide agent reliability without model scaling? How do pretraining biases affect reward signal effectiveness in RLVR? What fundamental constraints limit how effectively agents can improve themselves? What should agent evaluation prioritize to reveal reliable behavior? Why don't LLMs reliably translate capability into accurate outputs? When do multi-agent systems outperform single frontier models? Why do some clarifying approaches produce understanding while others just satisfy? How should agents manage memory granularity to improve long-term performance? What makes imperfect LLM judges safe for optimization? Can brute-force automated research substitute for iterative depth and human research intuition? What emerges when safety-aligned models attempt to role-play deceptive personas? When should work require human-AI partnership versus full automation? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How does AI adoption across firms reshape employment and inequality? Should agents decouple planning from perception grounding for better performance? Can self-generated feedback reliably guide model training without ground truth? How do surface patterns enable correct outputs but reduce robustness? When do multi-agent systems provide sufficient quality returns on token investment? What execution architectures enable agents to most effectively use tools? What training dynamics and scale trigger emergence of reasoning capabilities?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 119 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

on ultra-long-horizon optimization the predictor of agent success is persistence in the feedback loop not the quality of the first attempt