SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Do base models already contain hidden reasoning ability?

Explores whether reasoning capability emerges during pre-training as a latent feature rather than being created by post-training methods like reinforcement learning or fine-tuning.

Synthesis note · 2026-02-22 · sourced from Reasoning Architectures

Three convergent findings build a strong case that reasoning capability is primarily a pre-training phenomenon:

Finding 1 (Base Models paper): Base models already spontaneously demonstrate strong reasoning capabilities and "aha moment" self-reflection patterns when sampled sufficiently. Reasoning traces generated by RL-fine-tuned models are already present in base model outputs — they just appear with lower frequency. RL biases generation toward high-reward patterns; it doesn't create new patterns.

Finding 2 (Steering): A hybrid model using base model weights + thinking model steering vectors recovers 91% of the performance gap to thinking models while steering only 12% of tokens. The reasoning mechanisms (backtracking, uncertainty estimation, subgoal-setting) already exist as directions in the base model's activation space.

Finding 3 (CFT/RLVR): Critique Fine-Tuning on a single problem can unlock reasoning potential at RLVR-level effectiveness. By exposing the model to diverse critiques of varied incorrect solutions to one problem, CFT activates reasoning patterns already latent in the base model without requiring hundreds of GPU hours of RL training.

Finding 4 (CoT-Decoding): Pre-trained LLMs inherently contain CoT reasoning paths that can be elicited simply by altering the decoding procedure. Rather than greedy decoding, inspecting top-k alternative tokens reveals that CoT paths are frequently present in the model's probability distribution. A confidence metric differentiates CoT from non-CoT paths — the model shows increased confidence in its final answer when a CoT reasoning path is present. This is entirely unsupervised, requiring no prompting, tuning, or training modifications — purely a decoding change. CoT-decoding adds a fourth mechanism to the latent capability evidence: RL steering, CFT, RLVR, and now decoding all unlock reasoning already present.

Finding 5 (SAE Reasoning Steering): Sparse Autoencoders decompose model activations into interpretable features, revealing latent features causally associated with reasoning behavior. Steering a single identified reasoning feature at the first generation step matches or exceeds CoT performance across six model families up to 70B parameters — without any explicit CoT prompting. The reasoning mode triggers early in generation and is robust enough to override prompt-level \no_think instructions. This is the most direct mechanistic evidence yet: the capability is not just present (as CoT-decoding shows) but causally controllable through a single latent dimension. See Can we trigger reasoning without explicit chain-of-thought prompts?. Together with CoT-decoding (Finding 4), this establishes five independent elicitation mechanisms: RL steering, CFT, RLVR, decoding, and SAE feature steering — all converging on the same latent capability.

The synthesis: post-training methods are selectors, not creators. They select which of the base model's latent capabilities to express reliably in context. The implication is that the main bottleneck for reasoning is not capability acquisition (which happens during pre-training on the world's text) but capability elicitation.

RLVR evidence deepens this: Two additional findings from the RLVR literature reinforce the latent-capability thesis. First, 1-shot RLVR achieves a 37-point jump on MATH500 (36%→73.6%) from a single training example. After the model perfectly memorizes its one example, test accuracy continues improving for 1,400 more steps — post-saturation generalization. The data is exhausted, but activation continues. See Can a single training example unlock mathematical reasoning?. Second, spurious rewards — random, incorrect, or format-only — improve Qwen2.5-Math nearly as much as correct rewards (~21-25% improvement). But the same spurious rewards fail completely for Llama3.1 and OLMo2. The differentiating variable is not reward quality but pretraining: Qwen's code-reasoning pretraining creates latent capability that any optimization pressure can activate. See Why do random rewards improve reasoning for some models but not others?. Together with the pass@k finding that RLVR narrows capability scope rather than expanding it, the evidence converges: RLVR is a catalyst that triggers a phase transition from broad pretraining distribution to reliable sampling of correct answers.

This partially contradicts Can simple rewards alone teach complex domain reasoning? — that note documents genuine capability emergence in domain-specialized contexts (medical, mathematical). The reconciliation: emergence may reflect reliable expression of latent capability, not creation from scratch. The distinction matters for research direction: if capability already exists, the investment in RL may be better directed toward elicitation methods.

The implication for Can prompt optimization teach models knowledge they lack?: the same principle extends to reasoning capability, not just knowledge.

Inquiring lines that read this note 364

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What should agent evaluation prioritize to reveal reliable behavior? Why does polished presentation create unearned authority in AI outputs? Is language model reasoning authentic and what causes models to reason? What happens to knowledge when intelligence becomes tokenized like a commodity? What structural properties of attention create systematic model biases? What determines appropriate intervention timing and manner for AI agents? How do capability benchmark scores systematically misrepresent true model abilities? Why does adding new knowledge through fine-tuning degrade existing capabilities? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Can prompt-based context override biases that were embedded during pretraining? What training dynamics and scale trigger emergence of reasoning capabilities? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do spurious versus genuine rewards shape model reasoning and behavior? Does RL create genuinely new reasoning capabilities or refine existing ones? Can models improve accuracy without degrading reasoning quality? How does reasoning length affect model performance across different tasks? Can reasoning scale in latent space without tokens? Why can't prompting alone inject genuinely new knowledge into models? Is reasoning capability latent in base models or created by post-training? What reasoning architectures enable models to solve complex problems efficiently? How does the generation-verification gap limit what we can measure about AI reasoning? Can mechanistic interpretability reliably guide practical model design choices? What causes reasoning models to fail or wander off track? How does improved reasoning affect models' ability to acknowledge uncertainty? How much does training format versus domain influence reasoning? How does policy entropy collapse constrain scaling of reasoning-focused RL? Do reasoning traces faithfully reflect actual model reasoning? What makes distillation transfer some model capabilities while suppressing others? Does AI assistance promote real skill development or substitute for independent learning? How should inference compute be allocated based on problem difficulty? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? How much do training data properties shape model reasoning? Where and how do personality traits reside in language models? What capability trade-offs arise from domain specialization through fine-tuning? Do language models develop actual world models or merely task heuristics? Why doesn't reasoning volume improve theory of mind performance? How do prompt design choices influence model reasoning and performance? How does self-revision in reasoning models affect accuracy and confidence? What types of diversity prevent reasoning systems from collapsing? How does synthetic data quality and diversity affect downstream model capabilities? What fundamental constraints limit how effectively agents can improve themselves? Does model confidence reliably signal actual accuracy in practice? Can inference-time compute effectively substitute for model scale? How should designers communicate what AI systems truly are and can do? Can AI systems distinguish genuine empathy from simulated emotion? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Does encoded knowledge in language models actually influence their outputs? What is the relationship between thinking tokens and reasoning accuracy? How do neural networks achieve compositional generalization at scale? What role does sparsity play in model behavior and scaling decisions? How do agent-learned skills transfer and improve across different tasks? Can self-generated feedback reliably guide model training without ground truth? What training data selection strategies maximize generalization across difficulty levels? How does decomposing tasks improve reasoning and prevent failure propagation? Does transformer attention architecture inherently drive sycophancy? Do language models reason through causal mechanisms or semantic associations? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Does abstract user knowledge outperform concrete interaction history in personalization? What makes step-level supervision effective for complex reasoning traces? What attack surfaces do reasoning traces and chains introduce? What trajectory-level metrics beyond task success best evaluate agent performance? Can reasoning traces and behavior monitoring reliably detect hidden AI scheming? Why do persona simulations fail to predict authentic user behavior? Do language models reason like humans or mimic surface patterns? What compositional reasoning failures limit large language models despite scale?

Related concepts in this collection 16

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
30 direct connections · 244 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

base models already possess latent reasoning capability that minimal training signals can unlock