SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Does RL post-training create reasoning or just deploy it?

Investigates whether reasoning capability emerges during RL fine-tuning or already exists in base models. Matters because it reshapes how we build and optimize reasoning systems.

Synthesis note · 2026-02-22 · sourced from Reasoning Architectures

Post angle — Medium/LinkedIn

The dominant story: DeepSeek R1, GPT-o1, and their successors acquire reasoning capability through RL post-training. RL teaches models to think step-by-step, to backtrack, to verify — capabilities they didn't have before.

The emerging counter-evidence is striking. A hybrid model using a base model's weights with a thinking model's deployment decisions — zero weight updates — recovers 91% of the performance gap to thinking models by steering only 12% of tokens. Base models already spontaneously produce reasoning traces identical to thinking model traces when sampled sufficiently. Single-problem CFT achieves RLVR-level reasoning gains. Activation-space vectors encoding "backtracking" and "uncertainty estimation" already exist in base model hidden states before any RL.

The reframe: pre-training is when reasoning capability is acquired; RL post-training teaches when to deploy it.

This is not a trivial distinction. "When" training is cheaper, less data-hungry, and less fragile than "how" training. If capability already exists, elicitation methods (structured tool-calling, steering vectors, targeted fine-tuning on single problems) become much more attractive than full RL pipelines.

The hook for readers: "We've been crediting the locksmith for the key."

Connections: Does RL teach reasoning or just when to use it?, Do base models already contain hidden reasoning ability?, Can modular cognitive tools unlock reasoning without training?

Inquiring lines that read this note 168

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does polished presentation create unearned authority in AI outputs? Does RL create genuinely new reasoning capabilities or refine existing ones? Can reasoning scale in latent space without tokens? Why do stronger reasoning capabilities create tradeoffs with instruction following? Is reasoning capability latent in base models or created by post-training? What reasoning architectures enable models to solve complex problems efficiently? Can mechanistic interpretability reliably guide practical model design choices? How does improved reasoning affect models' ability to acknowledge uncertainty? How much does training format versus domain influence reasoning? What capability trade-offs arise from domain specialization through fine-tuning? Can models improve accuracy without degrading reasoning quality? How does policy entropy collapse constrain scaling of reasoning-focused RL? Should agents decouple planning from perception grounding for better performance? What training dynamics and scale trigger emergence of reasoning capabilities? How much do training data properties shape model reasoning? Do language models develop actual world models or merely task heuristics? Do reasoning traces faithfully reflect actual model reasoning? Why can't prompting alone inject genuinely new knowledge into models? Can prompt-based context override biases that were embedded during pretraining? What types of diversity prevent reasoning systems from collapsing? How do agent-learned skills transfer and improve across different tasks? What makes distillation transfer some model capabilities while suppressing others? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Is language model reasoning authentic and what causes models to reason? What causes reasoning models to fail or wander off track? What fundamental constraints limit how effectively agents can improve themselves? How does reasoning length affect model performance across different tasks? What design and behavioral factors drive false consciousness attribution to AI? Does encoded knowledge in language models actually influence their outputs? Can intelligent routing over smaller models outperform scaling a single large model? What makes step-level supervision effective for complex reasoning traces? How should agents manage memory granularity to improve long-term performance? How do neural networks achieve compositional generalization at scale? Why does adding new knowledge through fine-tuning degrade existing capabilities? How do spurious versus genuine rewards shape model reasoning and behavior?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 188 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

thinking models learn when not how — the case that rl post-training is a deployment optimizer not a capability creator