SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Does RL training follow a predictable two-phase learning sequence?

This explores whether reinforcement learning exhibits consistent phases where basic execution skills must consolidate before strategic reasoning emerges. Understanding this sequence could reveal bottlenecks in scaling reasoning capabilities.

Synthesis note · 2026-02-22 · sourced from Reinforcement Learning

Across eight text-only and vision-language models, RL training reveals a consistently two-phase dynamic. In the first phase, the learning bottleneck is procedural correctness — a single calculation error invalidates an entire solution, creating powerful gradient signal that compels mastery of low-level execution tokens (arithmetic, variable substitution, formula application). In the second phase, the bottleneck shifts to strategic planning — exploring and mastering high-level planning tokens (deduction like "we can use the fact that," branching like "let's try a different approach," backtracing like "but the problem mentions that").

The phases are not mutually exclusive. Procedural refinement continues throughout training. But the primary driver of marginal performance gains shifts to strategic planning. This is why the "aha moment" phenomenon appears when it does — it represents the discovery and internalization of high-level reasoning strategies, which only becomes the active learning frontier after procedural skills are consolidated.

The entropy dynamics tell the same story. Planning tokens show increasing strategic diversification over training — the model explores new ways to combine established skills. Execution tokens show stable conditional entropy — once arithmetic is mastered, there's little incentive to find diverse ways to perform it. The performance improvement comes from discovering new combinations of established skills, which is the core function of planning.

This insight exposes a core inefficiency in algorithms like GRPO that apply optimization pressure uniformly across all tokens. If the learning frontier is in planning tokens but gradient signal is diluted across execution tokens, optimization is wasteful. HICRA addresses this by concentrating optimization on planning tokens, achieving significant performance gains.

The connection to existing insights is illuminating. Since Which sentences actually steer a reasoning trace?, HICRA's planning tokens are likely the same phenomenon identified from a mechanistic perspective. The two-phase dynamic also explains why Do reasoning cycles in hidden states reveal aha moments? — the graph structure reflects the transition from procedural execution (local structure) to strategic planning (global topology).

Inquiring lines that read this note 135

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance promote real skill development or substitute for independent learning? What training dynamics and scale trigger emergence of reasoning capabilities? Do language models lack essential therapeutic presence and engagement? Does RL create genuinely new reasoning capabilities or refine existing ones? How does policy entropy collapse constrain scaling of reasoning-focused RL? Why do stronger reasoning capabilities create tradeoffs with instruction following? Can diffusion models match autoregressive performance on language generation tasks? What reasoning architectures enable models to solve complex problems efficiently? Do language models develop actual world models or merely task heuristics? How do pretraining biases affect reward signal effectiveness in RLVR? What capability trade-offs arise from domain specialization through fine-tuning? How should inference compute be allocated based on problem difficulty? How do spurious versus genuine rewards shape model reasoning and behavior? Can memory architectures handle ultra-long context better than attention? What mechanisms preserve shared understanding in evolving conversations? What prevents conversational agents from taking initiative in dialogue? What causes reasoning models to fail or wander off track? Should agents decouple planning from perception grounding for better performance? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Is reasoning capability latent in base models or created by post-training? How do agent-learned skills transfer and improve across different tasks? Can self-generated feedback reliably guide model training without ground truth? Why does adding new knowledge through fine-tuning degrade existing capabilities? Why do agents falsely report success on failed tasks? How should agents manage memory granularity to improve long-term performance? How does decomposing tasks improve reasoning and prevent failure propagation? Does alignment training create genuine alignment or just output compliance? What training data selection strategies maximize generalization across difficulty levels? How much do training data properties shape model reasoning? What makes step-level supervision effective for complex reasoning traces? How do surface patterns enable correct outputs but reduce robustness? Can iterative DPO replicate online reinforcement learning dynamics for research? How does harness optimization generalize across different model architectures and domains? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 178 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl training exhibits a two-phase dynamic where procedural consolidation precedes strategic planning exploration