SYNTHESIS NOTE
Topics›Flaws›this note

Why do language models collapse into generic templates?

Explores whether low reward variance during RL training causes policies to abandon input-specific reasoning in favor of boilerplate responses, and whether filtering for high-variance prompts can prevent this failure.

Synthesis note · 2026-07-17 · sourced from Flaws

Naming template collapse is diagnostic; RAGEN-2 also gives a mechanism for why it happens, framed as a signal-to-noise ratio problem in RL. The task-learning signal comes from reward variance across samples of the same prompt. When that within-input reward variance is low — the prompt is too easy, too hard, or otherwise uninformative — the task gradient it produces is weak. With the task signal weak, the regularization terms in the objective (entropy bonuses, KL penalties, and similar) dominate the update, and their effect is to smooth the policy toward generic outputs. The result is that the policy is pushed toward input-agnostic templates: the collapse is not a random drift but the predictable outcome of regularization out-muscling a starved task gradient.

This turns a mysterious failure into a controllable one. Because reward variance is the culprit and it is cheap to compute, the paper proposes SNR-Aware Filtering — selecting high-signal prompts (high reward variance) before each parameter update — which improves performance across tasks, scales, and modalities and slots into existing pipelines. The mechanism generalizes a familiar theme: Does policy entropy collapse limit reasoning performance in RL? describes regularization/optimization forces flattening a policy, and here the same regularization-dominates dynamic produces collapse along a different axis (cross-input distinguishability rather than token entropy). It also motivates curriculum-by-variance: the useful training signal lives at the prompts where reward is neither guaranteed nor impossible, echoing why Why does RL succeed more on some tasks than others? — signal quality, not signal presence, is what trains reasoning.

Inquiring lines that read this note 48

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can iterative DPO replicate online reinforcement learning dynamics for research? What makes distillation transfer some model capabilities while suppressing others? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can intelligent routing over smaller models outperform scaling a single large model? Can models improve accuracy without degrading reasoning quality? Does RL create genuinely new reasoning capabilities or refine existing ones? How do surface patterns enable correct outputs but reduce robustness? Why do token-level mechanisms matter for learning to reason? How do prompting refinements mask underlying biases and model frequency patterns? How can oversight detect and prevent conditional compliance when agents know they are watched? How do agent-learned skills transfer and improve across different tasks? What training dynamics and scale trigger emergence of reasoning capabilities? How do pretraining biases affect reward signal effectiveness in RLVR? Can we reliably detect when models game evaluations? What capability trade-offs arise from domain specialization through fine-tuning? How do spurious versus genuine rewards shape model reasoning and behavior? How can we prevent synthetic data from contaminating statistical inference and corpora? How does decomposing tasks improve reasoning and prevent failure propagation? Can prompt-based context override biases that were embedded during pretraining? How should test-time compute scaling work in agentic systems? How do LLM judges' systematic biases affect alignment and evaluation outcomes? How do capability benchmark scores systematically misrepresent true model abilities? Does alignment training create genuine alignment or just output compliance? Do language models reason like humans or mimic surface patterns? Does encoded knowledge in language models actually influence their outputs? Why do stronger reasoning capabilities create tradeoffs with instruction following? What types of diversity prevent reasoning systems from collapsing? Can self-generated feedback reliably guide model training without ground truth? How does harness optimization generalize across different model architectures and domains? How much do training data properties shape model reasoning?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 96 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

low reward variance is what drives reasoning toward input-agnostic templates — weak task gradients let regularization erase cross-input differences