SYNTHESIS NOTE
Topics›Self Refinement Self Consistency Feedback›this note

Can language models improve themselves without any external training data?

Explores whether two language models playing against each other—one generating questions, one solving them—can create a self-improving loop. Matters because it would eliminate dependence on human-labeled datasets.

Synthesis note · 2026-02-22 · sourced from Self Refinement Self Consistency Feedback

Self-Questioning Language Models (SQLM) adapts asymmetric self-play from robotic manipulation (OpenAI, 2021) to language domains. Two RL agents: a proposer and a solver. Given only a topic specification (e.g., "algebra word problems"), the proposer generates questions and the solver attempts answers.

The reward structure creates natural difficulty calibration: the proposer is rewarded when problems are neither too easy nor too hard — punished for trivially solvable questions and for impossible ones. The solver is rewarded based on majority voting (sampling multiple solutions and checking consensus), serving as a proxy for correctness without ground-truth labels. For coding tasks, the proposer can generate unit tests, providing direct verifiability.

This creates an automatically calibrated curriculum. The proposer explores the space of possible problems at the frontier of the solver's capability — hard enough to be informative, not so hard as to produce only noise. As the solver improves, the proposer must generate harder problems to maintain its own reward, creating escalating difficulty without human intervention.

The mechanism addresses two fundamental limitations of self-improvement: (a) the need for external training data (the proposer generates all training problems) and (b) the need for external verification (majority voting provides approximate correctness). Both solutions are intrinsic — no human labels, no external reward models, no ground-truth answers.

The key risk inherits from Does self-consistency reliably reward correct answers during training? — the solver's majority-voting reward is the same proxy signal, vulnerable to the same reward hacking. But the proposer provides a natural counterforce: it actively searches for the solver's weaknesses, potentially surfacing problems where majority voting is miscalibrated.

The connection to intrinsic motivation research is direct — curiosity-driven exploration (prediction error, state entropy, Go-Explore) provides the theoretical foundation for why generating novel challenges produces better learning than rehearsing known solutions.

Inquiring lines that read this note 38

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we prevent synthetic data from contaminating statistical inference and corpora? How does self-revision in reasoning models affect accuracy and confidence? Can self-generated feedback reliably guide model training without ground truth? How do surface patterns enable correct outputs but reduce robustness? What makes distillation transfer some model capabilities while suppressing others? What training data selection strategies maximize generalization across difficulty levels? Can prompt-based context override biases that were embedded during pretraining? Do language models learn genuine understanding or just surface patterns? Can brute-force automated research substitute for iterative depth and human research intuition? What compositional reasoning failures limit large language models despite scale? What capability trade-offs arise from domain specialization through fine-tuning? Do language models possess genuine introspective self-awareness or only behavioral mimicry? What fundamental constraints limit how effectively agents can improve themselves? How does reasoning length affect model performance across different tasks? How does improved reasoning affect models' ability to acknowledge uncertainty? How does the generation-verification gap limit what we can measure about AI reasoning? How does harness optimization generalize across different model architectures and domains?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 133 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

asymmetric self-play enables self-improvement without external data by training a proposer to generate challenging questions for a solver