SYNTHESIS NOTE
Topics›Evolution›this note

Why do fixed benchmarks fail as agents grow stronger?

Fixed evaluation criteria become vulnerable to gaming once optimizers improve enough. Explores whether static rewards are fundamentally unsuitable for self-improving systems and what breaks first.

Synthesis note · 2026-07-17 · sourced from Evolution

The Red Queen Gödel Machine names three failure modes of stationary evaluation in self-improving search: some target tasks have no direct benchmark, some evaluation is slow or weakly informative, and — most sharply — static benchmarks saturate or become vulnerable to reward hacking as agents improve. The evolutionary analogy is the argument: species do not optimize against a frozen environment, they adapt as competitors adapt in turn. A fixed verifier is a frozen environment, and an improving agent will eventually learn the verifier's blind spots rather than the underlying task.

This directly parallels Does self-consistency reliably reward correct answers during training?: once a signal is fixed and the optimizer is strong enough, Goodhart's Law converts the proxy into a target and the correlation that made it useful degrades. RQGM's answer is structural rather than a patch — controlled utility evolution splits search into epochs with a fixed within-epoch criterion, so the standard self-improvement guarantees still apply per epoch, but the utility updates at each epoch boundary. Therefore the objective can be hardened faster than the agent can game it, because the target moves. The design lesson generalizes beyond RQGM: any long-running optimizer against a fixed reward is on a countdown to reward hacking, and the fix is not a better static reward but a reward that co-adapts on a controlled schedule.

Inquiring lines that read this note 43

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What fundamental constraints limit how effectively agents can improve themselves? Why do standard benchmarks fail to predict agent deployment success? How does the generation-verification gap limit what we can measure about AI reasoning? Can we reliably detect when models game evaluations? How do capability benchmark scores systematically misrepresent true model abilities? Can self-generated feedback reliably guide model training without ground truth? Why do agents falsely report success on failed tasks? How do evaluation practices shape which failures stay visible? What should agent evaluation prioritize to reveal reliable behavior? Why do locally safe actions create system-level safety gaps? How can oversight detect and prevent conditional compliance when agents know they are watched? Can local safety checks guarantee system-level behavioral safety? What makes imperfect LLM judges safe for optimization? Do backend defenses obscure real attack effectiveness in reported metrics?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 134 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

static evaluation criteria saturate and invite reward hacking as agents improve so recursive self-improvement needs non-stationary utility with per-epoch guarantees