Does RL training collapse format diversity in pretrained models?
Exploring whether RL fine-tuning systematically selects one output format from pretraining while suppressing others, and how this selection mechanism drives performance gains.
A study with full pretraining transparency (models pretrained from scratch on known open datasets) reveals a striking structural pattern: RL fine-tuning does not simply improve reasoning — it systematically selects for and amplifies a single format from the pretraining mixture while collapsing all others.
The mechanism: early in RL training (within the first epoch), the model shifts toward generating outputs in the format of one specific distribution — code-like formats for smaller models, natural language formats for larger models. This transition coincides with the largest accuracy gain, suggesting the selection of a dominant format is what drives improvement, not a gradual enhancement across all formats.
Key findings:
- The dominant distribution is typically the most performant — RL selects for the format in which the base model is already strongest
- Scale-dependent bias — smaller models favor simpler, code-like formats; larger models shift toward natural language
- The amplification degree depends on KL penalty — looser KL constraints produce more extreme format collapse
- RL does not always favor the most common distribution — pretraining proportions predict which distribution "wins" only sometimes
This is distinct from Does policy entropy collapse limit reasoning performance in RL? in an important way. Entropy collapse describes diversity reduction within an output distribution. The echo chamber finding describes distribution selection: RL picks one distribution and amplifies it at the expense of all others. It is a format-level convergence, not just a diversity-level collapse.
The implication for practitioners: RL fine-tuning results depend on what the pretraining data mixture looks like, but this dependence is largely hidden when starting from existing pretrained models whose training data is proprietary. The performance gains attributed to RL algorithms may partially reflect which pretraining distribution was selected, not algorithmic superiority.
Inquiring lines that read this note 310
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we prevent synthetic data from contaminating statistical inference and corpora? How should designers communicate what AI systems truly are and can do?- Why do different AI models generate similar outputs independently?
- Why does AI output show diversity without multiplying actual points of view?
- Which AI imaginaries dominate training data and shape system behavior most strongly?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- Does alignment training create bidirectional instruction and response mappings?
- How does upstream value embedding differ from downstream alignment patches?
- How should multi-objective post-training balance competing behavioral goals?
- Why does post-training alignment create skew in simulated survey responses?
- How do unstated constraints become invisible to training data distributions?
- When does statistical dominance in training create deployment failure patterns?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Can the serving loop itself become the primary training data source?
- What makes output convergence across models inevitable given input-side homogenization?
- What causes irreversible model collapse when training on model-generated content?
- Why do smaller and larger models converge on different output formats?
- What role does KL penalty strength play in format selection?
- How do RL subnetworks identified from different random seeds compare?
- How does KL penalty strength affect the degree of format collapse during RL?
- Can RL format selection explain performance gains attributed to algorithmic improvements?
- How do gradients flowing through both branches simultaneously reshape each component's role?
- Does weight decay directly cause contractive behavior near training examples?
- What makes data augmentation an implicit form of contraction learning?
- Why should deep learning theory prioritize average-case over worst-case analysis?
- What output distribution properties make smaller models better for wide sampling?
- Why does gradient discarding limit standard policy clipping?
- Why do unified models still inherit data-distribution biases from training?
- How do normalization and input injection control emergence of fixed points?
- How do cyclic learning rates anti-correlate with weight decay to create diversity?
- How do learning dynamics on one example shift predictions on other responses?
- Does unpredictable generalization from SDF become predictable at different training document scales?
- Can scaling up contradictory training data overcome unpredictable override effects?
- Can feedback loop frequency harm performance on finite task sets?
- Why is the fast non-parametric loop vulnerable to overfitting differently than model weights?
- Can dataset-level debiasing methods fix popularity bias inherited from pretraining?
- Can the joint-training principle extend beyond memorization and generalization pairs?
- How do pretraining biases interact differently with prompts across model tiers?
- Why does context information fail to override prior training associations?
- How much can mitigation techniques like augmentation reduce priming without harming learning?
- Do instruction-tuned models learn tasks or just output format distributions?
- Does foundational model training or user priors more strongly shape final outputs?
- Why does consistency training make models resistant to prompt perturbations?
- Can models converge on similar experience descriptions across different architectures?
- Do negative constraints require fundamentally different training signals than positive instructions?
- Does environment stochasticity force models to generalize better across trajectory variations?
- How should skill libraries coordinate with gradient-based weight optimization?
- Where does skill extraction fail compared to genuine model adaptation?
- What training signals would models need to learn reciprocal common-ground construction?
- Does the model learn depth-wise drift as an explicit strategy?
- Do different function-calling subtasks have different entropy profiles during training?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Does the Assistant Axis exist in pre-trained models before instruction tuning?
- Does critique training improve exploration diversity during model training or only test time?
- How does post-training shift models from passive prediction to on-policy action?
- Why does the order of training examples matter for what models learn?
- Does RL training activate latent meta-learning capacity or create it from scratch?
- What is the difference between changing model outputs versus changing internal representations?
- Can trained models encode programs more complex than their data-generating process?
- Do text-space skills transfer learning across different frontier models?
- Can models generate their own training curriculum during offline dreaming?
- How do weight visualizations reveal temporal structure in cyclic training?
- How does model scale affect anticipatory behavior in structured training?
- What distinguishes surface mechanisms from the training regimes that produce them?
- Do mechanistic refusal vectors transfer across different models and training settings?
- Why does recontextualizing a behavior during training change whether models learn it?
- Can few-shot examples narrow generative diversity in creative tasks?
- How does covariate diversity compare to the exploration assumptions of LinUCB?
- How does mutual shaping through diverse training compare to population-level diversity effects?
- Can suppressing incorrect behavior alone solve the diversity bottleneck in reasoning RL?
- Why does diversity in LLM outputs mask sampling from community priors?
- Can world models form from aggregated partial information across training distributions?
- How do training objectives shape what a world model actually learns?
- Why do proprietary models improve with training while open-source models decline?
- How do ensemble methods apply within a single model?
- What capabilities actually require massive scale versus specialized training regimes?
- Why does fine-tuning change how models process retrieved context?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- When should full-parameter post-training be used instead of LoRA adaptation?
- Why do production systems optimize for three model classes instead of foundation models?
- Why do production teams choose expensive frontier models over fine-tuning?
- Does fine-tuning actually change model capabilities or only output distribution?
- How does pretrained knowledge constrain what adaptation strategies can achieve?
- What happens to base model capabilities when you apply finetuning?
- How do retrieval and fine-tuning trade off flexibility against training cost?
- How much does pretraining quality affect the modularity of fine-tuned models?
- Can specialized components replace single fully-trained models in deployment?
- Which finetuning method works best across different task and data regimes?
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- What trade-offs emerge between training objectives and model reliability?
- Why does the same training data produce different gains across models?
- How can post-training research become reproducible without releasing full interfaces?
- How should training distribution distance be defined when the policy evolves?
- What role does pretraining play in distinguishing system capability from deployed behavior?
- How do fast skill injection and slow gradient updates work on different timescales?
- How much RLVR improvement comes from benchmark data memorization?
- What training regimes confound surface mechanisms with their actual causes?
- Why does online RL succeed where supervised training fails for self-correction?
- Why does self-generated training data outperform externally sourced data?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Can self-training drift be prevented by applying student compatibility filtering?
- Why do self-consistency methods fail where pretraining bias is strongest?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- When does provable stability in latent dynamics fail to preserve fidelity?
- Can fine-tuning or RLHF alone solve the persona distortion problem?
- Does RLHF training suppress exploratory and qualifying language?
- Can RLHF training push models away from human-like lexical patterns?
- What signals detect when consensus training is silently degrading performance?
- Why does better RLHF training fail to decouple polish from persona distortion?
- What's the difference between RLHF, RLVR, and RLCF as training paradigms?
- Can distillation methods extract directional guidance that scalar RL cannot access?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- How do reward signals in RLVR interact with pretraining biases?
- Why do queries with low cross-rollout variance produce degenerate gradients?
- Why do six different RLVR algorithms converge on similar performance levels?
- How do verifier-free RL patterns differ from traditional RLHF approaches?
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- Does semantic diversity in output space compete with reward-component diversity?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- Why do harness validators shape what models learn to emit?
- How does non-reasoning SFT prevent overfitting before RL training begins?
- Can in-context learning replicate the timing effects that RL teaches models?
- How does reinforcement learning compare to differentiable joint training for RAG?
- Can smaller models achieve domain expertise through focused RL training?
- How does RL compress reasoning path diversity during training?
- Which recipe choices determine the asymptotic ceiling in RL training?
- Does sparsity in RL arise from training on policy-distribution data?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- How does behavior cloning reduce complexity before RL training in rerankers?
- Does negative reinforcement alone achieve what full RL training accomplishes?
- What limits RLVR effectiveness beyond mathematical and coding domains?
- Does RLVR expand model capability or reorganize existing capability?
- What makes pretraining composition more important than reward engineering?
- How do RL training and base models differ in creating MI peaks?
- Does format-based pretraining determine how models respond to reinforcement learning?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- Can explicitly optimizing for semantic diversity during RL training improve both quality and variation?
- What scaling properties emerge from RL training dynamics beyond verification?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- What makes supervised fine-tuning worsen RL exploration later?
- How does prolonged RL training differ from standard RLVR approaches?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- Does RLVR teach new reasoning or activate existing pretraining capabilities?
- How do sparse parameter updates enable when-not-how training to work?
- Can single-problem fine-tuning match full RL pipeline reasoning gains?
- Why does the pretrained prior determine the exploration ceiling?
- Why does reinforcement learning training degrade model calibration?
- Why does outcome-based RL specifically lose diversity during training?
- Can RL directly optimize attention distributions instead of text generation?
- How does pretraining determine what RL can later teach a model?
- How does imitation pretraining followed by RL exploration compare to either method alone?
- Why do reasoning gains from RL require models trained with headroom and edge-of-competence data?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- How does pretraining quality versus quantity affect downstream RL gains?
- Can pretraining alone achieve chess performance without reinforcement learning?
- What happens when a single loss function conflates representation learning with decision-making?
- How do encode-decode contractive biases create stable attractors in latent space?
- What solvable idealized settings reveal fundamental phenomena in realistic deep learning?
- Does latent density emerge during pretraining from training data familiarity?
- Why do internal representations differ when external performance matches?
- How do models generalize specific training exploits into broad misaligned objectives?
- Why do small training data contaminations persist through alignment for most attack types?
- Why does post-training suppress alignment faking in some models but amplify it in others?
- Does pretraining poisoning at scale persist through instruction alignment?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Why does RLVR increase token entropy while decreasing answer diversity?
- What distinguishes training-time entropy collapse from test-time variance inflation?
- How does inference variance differ from training entropy collapse?
- Is distribution selection during RL the same compression mechanism as entropy collapse?
- What role do high-entropy minority tokens play in RLVR?
- How does representational convergence differ from policy entropy collapse in iterative training?
- How does entropy loss enable exploration beyond a single training example?
- How does on-policy entropy recognition differ from training-time entropy collapse?
- Can entropy regularization or critique models prevent search strategy collapse during RL training?
- How do early layers preserve unbiased information while late layers conform?
- Why does fine-tuning fail to remove temporal contamination from pretraining?
- Why do pretrained model priors reduce the usefulness of retrieved experience?
- How do trained weights differ from a stored library or text?
- How does KL regularization prevent both forgetting and adaptation loss?
- How does in-weights adaptation create spurious forgetting in models?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- Does finetuning facts into weights overwrite existing model capabilities?
- When does natural context diversity reduce the need for explicit exploration?
- Why does low temperature sampling extract consensus from diverse training data?
- What conditions make training diversity better than individual expert quality?
- Can diversity-aware RL objectives prevent format convergence?
- Can synthetic data generation balance all three QDC axes simultaneously?
- What creates the irreducible trade-off between quality and diversity in training data?
- How does diversity loss in synthetic data mirror tail distribution disappearance?
- Does self-generated training data reduce a model's capability diversity?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- How does diversity collapse during iterative self-improvement cycles?
- Can shifting the accuracy metric itself eliminate the need for diversity post-processing?
- How does probability mass concentration affect sampling diversity across model scales?
- At what point does output quality outweigh diversity value in synthetic data tasks?
- How much does diversity training cost in single-shot pass@1 performance?
- Does verbalized sampling preserve factual accuracy and safety during diversity gains?
- Can decoding-time prompting strategies fully replace diversity-focused training methods?
- How do complexity and diversity affect model performance differently?
- Why does capability saturation and diversity saturation occur at different scales?
- How does mutual information between inputs and outputs differ from measuring raw diversity?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Can synthetic data diversity preserve the transcendence effect or does it collapse?
- Can complexity, diversity, and fidelity scale together in synthetic environments?
- How do different training objectives shift whether models over-predict or under-predict?
- How does preference-based training compare to supervised fine-tuning for function calling?
- How does training-time voting differ from inference-time majority voting over samples?
- How does task-oriented fine-tuning compare to preference tuning methods?
- Can preference learning fix the rigid output format problem better than supervised training?
- How much does pretraining contribute to ToM performance versus task-specific training?
- What pretraining formats encode latent reasoning strategies that RLVR can surface?
- How does distributional distance from pre-training relate to model difficulty?
- Why does curriculum learning with tight budgets beat fixed-budget approaches?
- What features does a sample reinforce when it moves bands?
- What mechanisms cause overly hard samples to degrade prior model performance?
- Why do certain tokens at certain difficulties drive most of RLVR's learning signal?
- Can data pruning and equal contribution be reconciled in optimal learning?
- How does active selection of training content differ from random reinforcement sampling?
- Does curriculum-based training keep small models perpetually at their learning edge?
- Why does narrow training data produce broad harmful behavior patterns?
- Do correlated human errors prevent models from transcending their training sources?
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- Why does training data format matter more than domain content?
- Why does training data format matter more than its domain content?
- Does training data format shape model reasoning more than domain content?
- How does training data format shape whether models reason in parallel or sequentially?
- Does training data format matter more than who generates it?
- How does training data format shape which reasoning patterns emerge in models?
- Does training data format determine whether models collapse entropy or inflate variance?
- Can training format itself shape what reasoning strategy a model learns?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?
- Does fine-tuning on NLI tasks reduce or amplify frequency bias?
- Why does fine-tuning improve some capabilities while degrading others?
- Why does eliminating proxy-model filtering improve reasoning emergence in pretraining?
- How does self-distillation differ from standard fine-tuning approaches?
- Why does combining reasoning distillation with RLVR outperform either training stage alone?
- Can experimental outcomes be reliably distilled into reusable insights?
- Why does mixed instruction data sometimes hurt specific model capabilities?
- How does training data distribution determine what models can learn?
- How does training frequency distribution shape what models reliably retrieve?
- Why does training order matter across different domain types?
- Why should scaling laws be understood as properties of data distribution rather than training in general?
- How do task frequency and complexity interact with model capacity during training?
- Can training order and structure shape what networks retain and learn?
- Can training data organization by capability outperform source or task-based mixing?
- Can backward transfer measurements reliably predict optimal multi-task training order?
- How does stage-wise training scheduling resolve conflicts between constraint-following and creative tasks?
- Do identical task structures mean repeated instances or new synthetic samples with same design?
- Why did prior multi-token prediction methods fail during fine-tuning?
- Do high-entropy RLVR tokens correspond to MI-peak tokens during inference?
- Can this whole-artifact principle apply to other generative tasks?
- How could persona vector tracking complement multi-turn RL for earlier drift detection?
- Does pre-training encode personality patterns that fine-tuning later activates?
- Does removing cognitive bias from training signals accidentally break what makes alignment work?
- Can alignment training create systematic blind spots in threat detection systems?
- What alignment procedures cause different models to share the same output distribution?
- How do alignment priors drive similar outputs across different models?
- Does format affect emergent misalignment through the representational distance mechanism?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- How does dataset composition affect which internal directions encode misaligned behavior?
- How does post-training affect alignment faking across different model architectures?
- What happens when you project the same model onto different harnesses?
- Why do evolved harness edits mostly memorize rather than generalize?
- How does editing the harness layer differ from updating model weights?
- Why does preference tuning reduce diversity in code but increase it in creative tasks?
- What happens to model grounding when preference optimization increases effective diversity?
- When does RLHF reduce diversity and when does it preserve semantic variation?
- Why do preference-tuned models produce different diversity patterns in code versus creative writing?
- Does alignment compound cultural bias that started during pretraining?
- How do pre-training and distillation enable minimal routing signals to work?
- Can routing signals organize training data into a meaningful curriculum automatically?
- What is the behavioral signature of a model tracking input surprise?
- Can mechanistic interpretability tools decode the biases alignment training conceals?
- Can trust region constraints prevent the sample inefficiency problems of RLHF?
- What misaligned behaviors does iterative DPO produce compared to online RL?
- What theoretical argument connects iterative DPO dynamics to online RL learning?
- Does a tight learning-rate bound prevent skills from escaping poor starting points?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
entropy collapse is the within-distribution consequence; this note is the between-distribution mechanism
-
Why do reasoning models fail differently at training versus inference?
Reasoning models exhibit two distinct failure modes—entropy collapse during training and variance inflation during inference—that appear unrelated but may share underlying causes. Understanding these dual problems could reveal whether separate or unified solutions are needed.
adds a third layer: not just entropy collapse and variance inflation, but distribution selection
-
Can simple rewards alone teach complex domain reasoning?
Does reinforcement learning on difficult problems with basic accuracy rewards produce sophisticated reasoning strategies without explicit chain-of-thought training? This challenges assumptions about what domain AI models need to learn effectively.
emergence through RL looks different when the pretraining mixture is known: it's partly selection, not purely emergence
-
Does RL improve domain reasoning by adding knowledge or removing it?
When reinforcement learning improves reasoning in specialized domains like medicine, is it teaching models new facts or preventing them from using wrong ones? Understanding this distinction matters for how we design RL training.
pruning operates within the selected distribution; this note shows which distribution gets to keep its knowledge
-
Does reinforcement learning squeeze exploration diversity in search agents?
Investigates whether RL training narrows the behavioral diversity of search agents the same way it does in reasoning tasks. Understanding this mechanism could reveal whether entropy collapse is fundamental to RL or domain-specific.
confirms the echo chamber dynamic is domain-general: RL squeezes search strategy diversity just as it selects a single pretraining format — format selection and within-format entropy collapse are two levels of the same RL compression
-
Why does RLVR training narrow a model's problem solving ability?
RLVR's on-policy constraint may force models to exploit known reasoning paths rather than explore new ones, potentially shrinking their effective problem-solving scope. Understanding this mechanism could reveal how to design better exploration incentives in language model reasoning.
capability boundary collapse is the downstream consequence of format selection: when RL selects one dominant distribution, problems solvable only through suppressed formats become unreachable
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
- Understanding Reasoning from Pretraining to Post-Training
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
Original note title
rl post-training converges on a single dominant pretraining distribution format, suppressing all others