SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can AI systems improve themselves through trial and error?

Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.

Synthesis note · 2026-02-23 · sourced from Novel Architectures

The original Gödel Machine proposed self-improving AI via provably beneficial self-modifications. In practice, formally proving the impact of most self-modifications is impossible. The Darwin Gödel Machine (DGM) replaces formal proofs with empirical validation: try modifications, test them on benchmarks, keep what works. This mirrors biological evolution — mutations are not verified in advance but produced, trialed, and selected.

DGM alternates between self-modification and evaluation phases. During self-modification, agents from the archive generate modified versions of themselves — rewriting their own code. During evaluation, each modified agent is tested on coding benchmarks. The key assumption: improvement on coding benchmarks indicates better coding capabilities, which in turn indicates better ability to self-modify. This creates a meta-competence loop: better coding → better self-modification → better coding.

Results: SWE-bench from 20.0% to 50.0%, Polyglot from 14.2% to 30.7%.

The evolutionary archive is critical. Inspired by open-endedness research, DGM maintains a growing library of all generated agent variants — including suboptimal but interesting ones. These serve as stepping stones for future generations, enabling diverse exploration paths. The system doesn't just optimize for immediate performance; it accumulates diverse capabilities that may enable future breakthroughs. This is fundamentally different from single-trajectory self-improvement.

Concrete improvements discovered include better code editing tools, long-context window management, and peer-review mechanisms — capabilities the original agent lacked that emerged through the self-improvement process.

The Python-based implementation makes the self-modification space Turing-complete in principle. The current version modifies agent design (tools, workflows) with frozen foundation models. Full self-improvement — rewriting training scripts, training new foundation models — is left as future work.

This directly addresses What limits how much models can improve themselves?: DGM circumvents the formal proof requirement by using empirical validation, but inherits a different limitation — improvement is bounded by what the benchmark can measure. The archive approach partially addresses How quickly do errors compound during model self-training? by maintaining diverse populations rather than following single improvement trajectories.

Inquiring lines that read this note 133

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How should designers communicate what AI systems truly are and can do? Can local safety checks guarantee system-level behavioral safety? Why does polished presentation create unearned authority in AI outputs? Does encoded knowledge in language models actually influence their outputs? How do evaluation practices shape which failures stay visible? How does the generation-verification gap limit what we can measure about AI reasoning? What fundamental constraints limit how effectively agents can improve themselves? Can self-generated feedback reliably guide model training without ground truth? When do multi-agent systems outperform single frontier models? How can evolutionary algorithms maintain diversity during solution search? Can inference-time compute effectively substitute for model scale? Do reasoning traces faithfully reflect actual model reasoning? How does self-revision in reasoning models affect accuracy and confidence? Can prompt-based context override biases that were embedded during pretraining? What should agent evaluation prioritize to reveal reliable behavior? What training dynamics and scale trigger emergence of reasoning capabilities? How do standardized protocols improve multi-agent coordination and reliability? How does synthetic data quality and diversity affect downstream model capabilities? How do capability benchmark scores systematically misrepresent true model abilities? Can brute-force automated research substitute for iterative depth and human research intuition? How do agent-learned skills transfer and improve across different tasks? How should agent systems validate and persist generated code artifacts? What makes imperfect LLM judges safe for optimization? Why do agents falsely report success on failed tasks? How effectively can language models perform reasoning, especially combined with symbolic methods? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How do surface patterns enable correct outputs but reduce robustness? Does preference optimization systematically degrade conversational grounding in language models? Can harness architecture and protocols provide agent reliability without model scaling? How should inference compute be allocated based on problem difficulty? How does harness optimization generalize across different model architectures and domains? How can we prevent synthetic data from contaminating statistical inference and corpora? How does evaluation scope and dimensionality affect what we measure? Can models improve accuracy without degrading reasoning quality? How do pretraining biases affect reward signal effectiveness in RLVR? Do reasoning benchmarks predict model performance in long-horizon workflows? Do honeypot benchmarks validly measure reward hacking better than standard tests?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
22 direct connections · 177 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

darwin godel machine achieves open-ended self-improvement by replacing formal proofs with empirical validation and evolutionary archives