SYNTHESIS NOTE
Topics›Conversation Agents›this note

Does extended thinking help or hurt model reasoning?

Explores whether activating thinking mode improves reasoning performance, and what role training plays in determining whether extended internal reasoning chains are productive or counterproductive.

Synthesis note · 2026-02-22 · sourced from Conversation Agents

The proactive critical thinking experiments reveal a striking interaction between training and inference-time reasoning. For vanilla (off-the-shelf) models, activating "thinking mode" — the extended internal reasoning chains used by models like Qwen3 — actually degrades performance on proactive critical thinking tasks. The extended thinking "appears to induce counterproductive self-doubt rather than useful analysis, leading to a clear drop in performance."

But after RL training on proactive critical thinking tasks, the same thinking mode becomes beneficial. Training fundamentally changes how models use their internal reasoning. This is not merely about more or less thinking — it is about the quality direction of thinking.

The finding connects to several established insights but adds a distinct mechanism:

Since Does RL teach reasoning or just when to use it?, RL manages the timing of reasoning. The proactive thinking result extends this: RL also manages the mode of reasoning — redirecting extended thinking from unproductive self-doubt toward productive gap analysis.

The SFT finding adds nuance: when SFT data is self-generated by the model, it "does not inherently enhance its capabilities" and may reduce output entropy, constraining the subsequent RL phase. This echoes Does policy entropy collapse limit reasoning performance in RL? — SFT-then-RL may face the same entropy collapse that pure RL faces, but through a different mechanism (entropy reduction from self-generated imitation rather than RL convergence).

The practical implication: extended thinking is not a universal good. It is a resource that can be directed productively or destructively, and the direction depends on training. "More thinking" applied to a model without the right training signal may systematically make things worse.

Inquiring lines that read this note 134

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models reason like humans or mimic surface patterns? How do spurious versus genuine rewards shape model reasoning and behavior? Can models improve accuracy without degrading reasoning quality? How does reasoning length affect model performance across different tasks? Is reasoning capability latent in base models or created by post-training? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? How does improved reasoning affect models' ability to acknowledge uncertainty? Why do stronger reasoning capabilities create tradeoffs with instruction following? What is the relationship between thinking tokens and reasoning accuracy? Is language model reasoning authentic and what causes models to reason? What happens to knowledge when intelligence becomes tokenized like a commodity? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Does model confidence reliably signal actual accuracy in practice? Can reasoning scale in latent space without tokens? How does evaluation scope and dimensionality affect what we measure? What causes reasoning models to fail or wander off track? Do language models reason through causal mechanisms or semantic associations? Why doesn't reasoning volume improve theory of mind performance? How does persona conditioning amplify demographic stereotyping and bias in models? What structural distinctions matter in reasoning and argumentation? How does self-revision in reasoning models affect accuracy and confidence? Do reasoning traces faithfully reflect actual model reasoning? Does RL create genuinely new reasoning capabilities or refine existing ones? Can AI systems distinguish genuine empathy from simulated emotion? How do prompting refinements mask underlying biases and model frequency patterns? What training dynamics and scale trigger emergence of reasoning capabilities? Does transformer attention architecture inherently drive sycophancy? What makes step-level supervision effective for complex reasoning traces? Why does polished presentation create unearned authority in AI outputs?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 164 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl training transforms thinking mode from counterproductive self-doubt into beneficial proactive analysis — the same mechanism helps or hurts depending on training