SYNTHESIS NOTE
Topics›Evaluations›this note

Can dense subtask grading reveal agent progress on ultra-long tasks?

When agent tasks stretch to hours and hundreds of episodes, does breaking them into fine-grained graded subtasks expose meaningful progress that binary pass-fail scoring completely erases? This matters because most agents fail the final outcome anyway.

Synthesis note · 2026-07-17 · sourced from Evaluations

Long-Horizon-Terminal-Bench takes the standard Terminal-Bench setup — a reference solution or simulation engine per task — but decomposes each of its 46 tasks into fine-grained deterministic subtasks, each graded against the environment. This produces dense intermediate rewards and partial credit, which matters because the tasks are genuinely long: averaging 239 episodes, 9.8M tokens, and 88.9 minutes per run. On workflows that long, a binary final-outcome verdict discards almost all the information. The strongest model (Grok 4.5) passes only 28.3% at R≥0.95 while the mean across 17 models is just 6.4% — a near-floor result that outcome-only scoring would render as undifferentiated failure.

The design argument is that reward density is not a nicety but a measurement necessity once horizons stretch past one-shot problem solving. When agents make "meaningful but incomplete progress," a pass/fail benchmark cannot distinguish a model that got 80% of the way from one that stalled immediately — both read as zero. This grounds the broader claim that How should we measure agent system performance beyond task success?: dense subtask grading is one concrete instrument for that shift. It also echoes the training-side finding that Can we reward reasoning steps without human annotation? — dense signals expose structure that outcome-only aggregation flattens, whether the goal is evaluation or optimization.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What should agent evaluation prioritize to reveal reliable behavior? How can infrastructure records verify actual agent behavior? What trajectory-level metrics beyond task success best evaluate agent performance?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 151 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

dense subtask grading on long-horizon terminal tasks turns agent evaluation from a pass-fail verdict into a progress measurement