SYNTHESIS NOTE
Topics›Test Time Compute›this note

Do critique models improve diversity during training itself?

Explores whether critique integrated into the training loop, beyond test-time scoring, actively maintains solution diversity and prevents the model from converging too narrowly during iterative self-training.

Synthesis note · 2026-02-20 · sourced from Test Time Compute

The intuitive framing of critique models is that they help at test time: the model generates, the critic scores, we select the best. But the more important finding from AutoMathCritique is that critique integrated into the training loop improves the actor model's exploration efficiency and solution diversity during training itself.

Without critique in the loop, iterative self-training suffers from "tail narrowing" — the model converges on a narrow distribution of solutions, becoming less able to explore diverse reasoning paths. The critique model counteracts this: by providing step-level feedback on exploration, it guides the actor toward high-quality paths it wouldn't have discovered alone, maintaining distributional breadth through training.

This connects to Does policy entropy collapse limit reasoning performance in RL?: critique models are a way to maintain entropy — the exploration needed for continued improvement — without relying solely on architectural entropy management (Clip-Cov, KL-Cov). The critique is an external signal that prevents premature convergence.

The implication: critique models are training infrastructure as much as inference infrastructure. Evaluating them only on test-time accuracy misses their more fundamental role.

Inquiring lines that read this note 94

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? How do prompting refinements mask underlying biases and model frequency patterns? What types of diversity prevent reasoning systems from collapsing? How do spurious versus genuine rewards shape model reasoning and behavior? Is reasoning capability latent in base models or created by post-training? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can prompt-based context override biases that were embedded during pretraining? Can self-generated feedback reliably guide model training without ground truth? How should designers communicate what AI systems truly are and can do? How well do AI systems understand human social norms? How do pretraining biases affect reward signal effectiveness in RLVR? How does synthetic data quality and diversity affect downstream model capabilities? Do structural constraints outperform deep architectures in recommendation systems? What capability trade-offs arise from domain specialization through fine-tuning? What makes step-level supervision effective for complex reasoning traces? What training data selection strategies maximize generalization across difficulty levels? What fundamental constraints limit how effectively agents can improve themselves? How does evaluation scope and dimensionality affect what we measure? How does improved reasoning affect models' ability to acknowledge uncertainty? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? How does self-revision in reasoning models affect accuracy and confidence? What safeguards enable trustworthy AI-assisted scientific peer review at scale? Why don't LLMs reliably translate capability into accurate outputs? Why can't prompting alone inject genuinely new knowledge into models? What training dynamics and scale trigger emergence of reasoning capabilities? Does RL create genuinely new reasoning capabilities or refine existing ones? Does preference optimization systematically degrade conversational grounding in language models? How should inference compute be allocated based on problem difficulty? What causes reasoning models to fail or wander off track? How can evolutionary algorithms maintain diversity during solution search? How do surface patterns enable correct outputs but reduce robustness? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How does harness optimization generalize across different model architectures and domains? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 152 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

critique models improve exploration diversity during training not just test-time accuracy