Does extended thinking help or hurt model reasoning?
Explores whether activating thinking mode improves reasoning performance, and what role training plays in determining whether extended internal reasoning chains are productive or counterproductive.
The proactive critical thinking experiments reveal a striking interaction between training and inference-time reasoning. For vanilla (off-the-shelf) models, activating "thinking mode" — the extended internal reasoning chains used by models like Qwen3 — actually degrades performance on proactive critical thinking tasks. The extended thinking "appears to induce counterproductive self-doubt rather than useful analysis, leading to a clear drop in performance."
But after RL training on proactive critical thinking tasks, the same thinking mode becomes beneficial. Training fundamentally changes how models use their internal reasoning. This is not merely about more or less thinking — it is about the quality direction of thinking.
The finding connects to several established insights but adds a distinct mechanism:
Since Does RL teach reasoning or just when to use it?, RL manages the timing of reasoning. The proactive thinking result extends this: RL also manages the mode of reasoning — redirecting extended thinking from unproductive self-doubt toward productive gap analysis.
The SFT finding adds nuance: when SFT data is self-generated by the model, it "does not inherently enhance its capabilities" and may reduce output entropy, constraining the subsequent RL phase. This echoes Does policy entropy collapse limit reasoning performance in RL? — SFT-then-RL may face the same entropy collapse that pure RL faces, but through a different mechanism (entropy reduction from self-generated imitation rather than RL convergence).
The practical implication: extended thinking is not a universal good. It is a resource that can be directed productively or destructively, and the direction depends on training. "More thinking" applied to a model without the right training signal may systematically make things worse.
Inquiring lines that read this note 134
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do language models reason like humans or mimic surface patterns? How do spurious versus genuine rewards shape model reasoning and behavior?- Do spurious rewards activate reasoning without teaching new skills?
- How do reward models benefit from extended thinking during evaluation scoring?
- What role does task structure play in rewarding delayed thinking?
- What distinguishes genuine reasoning activation from memorization-assisted answer recall?
- Can benchmark improvements hide degradation of deliberative reasoning?
- When should action deliberation trigger during reasoning steps?
- Do explicit reasoning chains improve or harm performance on complex judgment tasks?
- Can extended thinking genuinely improve reasoning or just increase variance?
- Why do more capable models prefer shorter chains of thought?
- Does explicit reasoning help or hurt tasks requiring continuous nuanced judgment?
- How does difficulty level change whether extended thinking provides genuine reasoning signal?
- When does explicit reasoning actually degrade performance on a task?
- Can extended reasoning training capture individual strategic thinking styles?
- Why does step-by-step reasoning degrade performance on judgment-based tasks?
- Does distillation from reasoning models spread overthinking to smaller models?
- Why does extended thinking increase output variance without improving reasoning quality?
- Does deep-thinking ratio measure computational effort better than chain-of-thought length?
- Can extended deliberation in agents become counterproductive like human overthinking?
- Why does inference-time thinking hurt proactive critical thinking in vanilla models?
- Can models trained on longer contexts develop better fundamental reasoning abilities?
- Does explicit reasoning help or hurt tasks requiring continuous judgment?
- How does extended thinking affect variance in reasoning model outputs?
- When should a system choose extended thinking versus quick responses?
- How should timing for reasoning intervention be determined during inference?
- How much does extended thinking actually improve model reasoning ability?
- Does penalizing thought transitions improve reasoning without model retraining?
- Does more thinking always improve language model accuracy?
- Can a single model implement fast thinking, slow thinking, and tool use?
- Does performative reasoning mask underlying uncertainty even on easy problems?
- Why do longer reasoning chains explore like tourists instead of scientists?
- Why do thinking models execute longer tasks than standard language models?
- Is premature decision-making a form of underthinking in transformer models?
- Why does reflection in reasoning models often become theater rather than genuine thought?
- When does extended thinking hurt performance on easier problems?
- How does flip-event regression differ from premature thought path abandonment?
- Do earlier errors in long tasks increase the likelihood of future mistakes?
- Does longer reasoning always improve model accuracy on complex tasks?
- What makes training-free approaches like Soft Thinking preferable to SoftCoT?
- What makes reasoning capability a pre-training rather than post-training phenomenon?
- Can RL training teach models when to activate reasoning versus when to skip it?
- How do reasoning training methods sacrifice some thinking skills while improving others?
- Can activation-space steering vectors replicate thinking model performance without retraining?
- What other triggers can activate the latent reasoning capability?
- Does RL training actually restore the critical thinking that reasoning models lose?
- What is the distinction between teaching reasoning how versus when to activate?
- Can pretraining signals unlock latent reasoning that post-training merely activates?
- What distinguishes reasoning activation mechanisms across different training methods?
- Can thought quality alone be trusted to guide model training?
- How do timing and search internalization interact during reasoning post-training?
- Can structured questioning prompts improve reasoning beyond standard conversational training?
- Why does pre-training provide the raw material for emergent thinking?
- Can we predict when a model will develop thinking behaviors?
- Why does extended reasoning training improve exploration without adding new capabilities?
- Does targeting the edge of competence during RL pretraining unlock true reasoning gains?
- What makes some reasoning strategies genuinely novel versus latent?
- Do high-influence thoughts align with SAND deliberation triggers?
- What three factors actually drive chain of thought performance improvements?
- Does chain-of-thought reasoning help or hurt social reasoning tasks?
- What distinguishes metacognitive regulation from standard chain-of-thought reasoning?
- How do thought actions represent policy improvement steps in practice?
- Does chain-of-thought trigger latent reasoning or create it?
- How much did social chain-of-thought prompting improve each model family's strategic reasoning?
- Can proactive critical thinking alone enable models to request clarification effectively?
- Does training for better reasoning reduce an AI system's ability to abstain?
- Can proactive critical thinking train models to request clarification actively?
- How does proactive critical thinking enable models to identify missing information?
- Does reasoning training actively undermine the abstention capacity safety training created?
- Why must procedural skills consolidate before strategic reasoning can develop?
- Do reasoning models trade instruction following for deliberative capability?
- Why do human-curated thought examples fail to improve model thinking?
- Does reasoning structure match explicit versus implicit task demands?
- Can models distinguish between activated knowledge and genuine reasoning?
- Why does latent reasoning override no-think instructions in models?
- Do shorter reasoning chains maintain instruction adherence better than longer ones?
- Why do foundation models develop task-specific heuristics instead of causal understanding?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- Can instruction-level interventions fix memory-induced reasoning failures in practice?
- Can budget-tightening curricula improve reasoning efficiency more than fixed budgets?
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- What triggers overthinking versus underthinking in reasoning models?
- Why does reasoning accuracy degrade beyond a critical thinking token threshold?
- How do thinking tokens function as mutual information peaks in reasoning?
- What happens to reasoning accuracy when models use more thinking tokens?
- Does the thinking box provide genuine reasoning or just token budget?
- Can thinking token density explain reasoning performance beyond total length?
- Why do different model training approaches produce different overthinking thresholds?
- Does task difficulty alone determine how many thinking tokens a model should use?
- What happens to model reasoning accuracy as thinking token requirements exceed critical thresholds?
- Can conditioning generation on difficulty probes reduce overthinking on simple tasks?
- Can activation steering compress reasoning without retraining models?
- What causes reasoning accuracy to degrade beyond a critical thinking-token threshold?
- What accuracy gains come from adaptive versus fixed thinking budgets?
- Does a critical thinking token threshold exist for model accuracy?
- When should an LLM engage extended reasoning versus responding directly?
- Can extended thinking modes introduce genuine rhetorical exploration to LLMs?
- Can parallel thinking outperform sequential thinking under the same token budget?
- Why does parallel thinking outperform sequential thinking under fixed token budgets?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Does thought consolidation address the confirmatory reflection problem in reasoning models?
- How does collaboration itself become a degradation mechanism in reasoning tasks?
- Do reasoning failures stem from strategy or from calculation breakdown?
- Can models overthink and underthink at the same time?
- What causes reasoning quality to degrade during long research tasks?
- How does active reasoning through interaction differ from passive single-turn problem solving?
- How does o1-style reasoning relate to learned search processes versus memorized solutions?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- Why does reasoning effort fail to improve theory of mind performance?
- Does formal reasoning training actively degrade social reasoning ability?
- Do longer reasoning traces actually improve theory of mind accuracy?
- How do emotional and social simulations enable better hypothetical reasoning?
- Can reasoning scaffolds help with nuanced judgment tasks like empathy?
- Why might social reasoning work differently than formal logical reasoning?
- Why does reasoning volume fail to improve theory of mind performance?
- Does reasoning effort correlate with social reasoning accuracy?
- Why does revision often make reasoning accuracy worse in frontier models?
- Does internal self-revision actually degrade reasoning accuracy in models?
- Can extended RL training unlock genuinely new reasoning strategies models cannot discover otherwise?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does RL teach reasoning or just when to use it?
Does reinforcement learning in thinking models actually create new reasoning abilities, or does it simply teach existing capabilities when to activate? This matters for understanding where reasoning truly emerges.
RL manages timing; this paper shows RL also manages quality direction of reasoning
-
Can models learn when to think versus respond quickly?
Explores whether a single language model can adaptively choose between extended reasoning and direct responses based on task difficulty. This matters because it could make inference more efficient by allocating compute only when needed.
DeGRPO mode selection; proactive thinking adds a training-mediated quality dimension
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
SFT-then-RL may face entropy collapse through self-generated imitation
-
What critical thinking skills do reasoning models actually lose?
Step-by-step reasoning training optimizes narrow deductive thinking while degrading meta-cognitive abilities like recognizing futile thinking and maintaining tentative reasoning. Understanding this tradeoff matters for deploying reasoning models reliably.
the thinking-mode reversal is a specific instance of the broader critical thinking problem: reasoning training optimizes one narrow type of thinking while degrading others; the proactive thinking result shows RL can selectively repair one form of degradation (self-doubt → gap analysis) while the critical thinking post documents the broader pattern
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- Base Models Know How to Reason, Thinking Models Learn When
- Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- Rethinking Thinking Tokens: LLMs as Improvement Operators
Original note title
rl training transforms thinking mode from counterproductive self-doubt into beneficial proactive analysis — the same mechanism helps or hurts depending on training