SYNTHESIS NOTE
Topics›Reward Models›this note

Does training order reshape how models handle different task types?

Explores whether the sequence of multi-task RL training systematically affects model capabilities across structured and creative domains, and whether this ordering effect can be predicted and optimized.

Synthesis note · 2026-02-22 · sourced from Reward Models

The standard framing of Does policy entropy collapse limit reasoning performance in RL? treats entropy collapse as a uniform phenomenon — RL training decreases entropy. Omni-Thinker (2025) reveals this is domain-dependent: structured domains (math, coding) decrease output entropy, while open-ended domains (creative writing, dialogue) increase it.

This is not a minor observation — it makes training order a mechanistic variable, not just a scheduling convenience. If you train creative writing first and structured reasoning second, the structured training will collapse the entropy that creative training expanded, potentially degrading creative capability. If you train structured reasoning first and creative writing second, the creative training preserves and expands the model's expressive range. The ordering effect is predictable from backward transfer (BWT) measurements.

Omni-Thinker uses BWT-guided scheduling: order tasks so that later tasks experience minimal negative backward transfer from earlier tasks. The approach uses hybrid rewards — verifiable (rule-based) for deterministic domains + preference-based (LLM-as-Judge) for subjective domains — enabling unified training across domain types within a single policy. The "short-form" QA tasks condition on distractors to reduce reward hacking from random guessing.

The gains are substantial: 6.2% over joint multi-task training, 12.4% over model merging. The accuracy of final multi-task models is well-predicted by forgettability rankings, even under simplifying assumptions — suggesting BWT-guided scheduling has principled theoretical grounding.

This extends Does gradually tightening token budgets beat fixed budget training? from temporal budgets to task ordering: the dimension that matters for multi-task RL is not just how much compute per task, but which tasks come first. And it enriches the entropy collapse understanding: entropy collapse is not a bug to fix everywhere — in structured domains, it reflects desirable precision. The problem is when structured-domain entropy collapse propagates to damage open-ended capabilities.

Inquiring lines that read this note 104

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models develop actual world models or merely task heuristics? Why do stronger reasoning capabilities create tradeoffs with instruction following? Does RL create genuinely new reasoning capabilities or refine existing ones? What capability trade-offs arise from domain specialization through fine-tuning? How much do training data properties shape model reasoning? How does decomposing tasks improve reasoning and prevent failure propagation? How much does training format versus domain influence reasoning? How do prompting refinements mask underlying biases and model frequency patterns? How do prompt design choices influence model reasoning and performance? What training dynamics and scale trigger emergence of reasoning capabilities? Does AI assistance promote real skill development or substitute for independent learning? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can intelligent routing over smaller models outperform scaling a single large model? How does synthetic data quality and diversity affect downstream model capabilities? How do surface patterns enable correct outputs but reduce robustness? Can prompt-based context override biases that were embedded during pretraining? How do pretraining biases affect reward signal effectiveness in RLVR? What trajectory-level metrics beyond task success best evaluate agent performance? How does harness optimization generalize across different model architectures and domains? Can brute-force automated research substitute for iterative depth and human research intuition? Why does memory consolidation cause performance regression in continual learning? Why does adding new knowledge through fine-tuning degrade existing capabilities? Do language models lack essential therapeutic presence and engagement? Why do agents falsely report success on failed tasks? What training data selection strategies maximize generalization across difficulty levels? Does alignment training create genuine alignment or just output compliance? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Does model confidence reliably signal actual accuracy in practice? How do capability benchmark scores systematically misrepresent true model abilities? Can models improve accuracy without degrading reasoning quality? How do training data properties determine the emergence of internal misalignment? How should test-time compute scaling work in agentic systems? Do language models reason like humans or mimic surface patterns? How does evaluation scope and dimensionality affect what we measure?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 161 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

multi-task rl reveals complementary entropy dynamics — structured domains systematically decrease output entropy while creative domains increase it making training order a mechanistic variable