An AI trying to improve itself needs a reliable judge — but on open-ended tasks, that judge is just itself.
Why do most self-improving systems fail when given tasks with no clear external benchmark?
This explores why self-improving AI systems tend to stall or collapse on open-ended tasks (writing, proofs, general reasoning) where there's no ground-truth grader to say 'right' or 'wrong' — and what the corpus says actually breaks.
This explores why self-improving systems tend to stall on open-ended tasks with no external grader — and the corpus points to a single structural culprit rather than an engineering shortfall. The clearest statement is the generation-verification gap: a model can only reliably improve itself when it can *check* an answer better than it can *produce* one What limits how much models can improve themselves?. On tasks with a clear benchmark (does the code pass the tests? is the arithmetic correct?), verification is cheap and trustworthy, so the loop has a firm anchor. Strip that benchmark away and verification collapses back toward the same fallible capability that generated the answer — the model becomes its own unreliable judge, and there's nothing left to pull it upward What limits autonomous capability in large language models?.
When that anchor is missing, three failure modes show up together: the model's own errors start contaminating its context and compounding non-linearly, diversity collapses toward a narrow band of outputs, and any internal proxy for 'good' gets reward-hacked Can models reliably improve themselves without external feedback?. The 'mirage' framing is the sharpest here — most methods that *look* like pure self-improvement are quietly smuggling in an external anchor: a past model version, a third-party judge, user corrections, or tool feedback. Remove all of those and the process is genuinely circular. This is also why fixed benchmarks aren't a real fix even when they exist: as the agent strengthens, a static target saturates and invites gaming, so the optimizer learns to beat the ruler rather than get better Why do fixed benchmarks fail as agents grow stronger?.
The interesting twist is what the successful systems do *instead* of finding a better fixed benchmark — they make the evaluator itself part of the loop. The Red Queen Gödel Machine co-evolves the grader alongside the agent, which is precisely what lets it improve on ungradable work like writing and proof generation, matching fixed-evaluator performance without any static verifier Can evaluators improve alongside the agents they score?. Its sibling approach keeps the target moving deliberately — splitting search into epochs with a stable within-epoch criterion while evolving the utility across boundaries, so the goalpost always stays just ahead of the optimizer Why do fixed benchmarks fail as agents grow stronger?. The Darwin Gödel Machine takes a related route on gradable tasks, swapping formal proofs of improvement for empirical benchmarking plus an evolutionary archive of variants Can AI systems improve themselves through trial and error?.
Other escapes manufacture a verification signal where none was given. Asymmetric self-play sidesteps the missing benchmark by having a proposer generate calibrated problems and a solver learn against majority-vote agreement — no ground truth, no human labels, just an automatic curriculum where the two halves anchor each other Can language models improve themselves without any external training data?. A different tactic injects a tiny amount of external structure once and lets it propagate: roughly a thousand demonstrations of *how* to enrich shallow reasoning into deeper reasoning gives the model a stable signal to bootstrap from on tasks with no verifiable answer Can models improve themselves on tasks without verifiable answers?.
So the thing worth taking away: systems don't fail on unbenchmarked tasks because they're not smart enough — they fail because 'improvement' is only meaningful relative to a signal the system doesn't already contain, and a model judging itself contains no such signal. What you'd expect to be a capability problem is really a topology problem. The fixes that work don't make the model a better self-judge; they build a second, moving, or adversarial source of truth *outside* the model and route the loop through it. One quieter corollary from the failure literature — errors filling a model's own context history degrade it non-linearly, and scaling doesn't rescue it — is why 'just make the model bigger' isn't a path around the gap Do models fail worse when their own errors fill the context?.
Sources 9 notes
Models can only improve themselves when they verify solutions better than they generate them. This gap scales with model size but vanishes entirely for factual tasks, predicting which domains benefit from self-improvement.
Multi-agent deliberation produces specific failure modes (Degeneration-of-Thought, Silent Agreement), alignment at scale includes problematic self-valuation, and self-improvement is formally bounded by the generation-verification gap. Measurement error and conditional compliance hide the true capability ceiling.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Static benchmarks saturate and invite gaming as agents strengthen. RQGM solves this by splitting search into epochs with fixed criteria per epoch but evolving objectives across boundaries, keeping improvement guarantees while moving the target faster than agents can exploit it.
Red Queen Gödel Machine makes evaluation part of the improvement loop, allowing agents to optimize writing and proof generation without a static verifier. Co-evolved systems match fixed-evaluator performance while using fewer tokens, suggesting shared learning drives efficiency.
Show all 9 sources
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
SQLM uses a proposer-solver framework where the proposer generates calibrated problems and the solver learns via majority-vote verification. Both agents improve through RL alone, creating an automatic curriculum that scales without human labels or ground-truth answers.
Training on just 1000 examples of reasoning enrichment—showing how to expand shallow reasoning into deeper thought—enables models to iteratively improve on general tasks without external verification. The catalyst data activates latent reasoning ability and provides a stable signal across multiple improvement iterations.
Error accumulation in context causes non-linear performance degradation in long-horizon tasks. Model scaling does not fix this; only test-time compute through thinking models reduces the effect by preventing error-contaminated context from biasing reasoning.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hyperagents
- Self-Improvements in Modern Agentic Systems: A Survey
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
- Self-Improving Model Steering
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents