SYNTHESIS NOTE
Topics›Reasoning Critiques›this note

Does chain-of-thought reasoning actually generalize beyond training data?

Explores whether CoT's strong performance on benchmarks reflects genuine reasoning ability or merely reflects learned patterns tied to specific distributions. Tests how CoT behaves when tasks, formats, or reasoning length shift away from training data.

Synthesis note · 2026-02-22 · sourced from Reasoning Critiques

Chain-of-Thought prompting performs well on in-distribution problems and fails predictably as distributional discrepancy increases. This is not a bug — it is the fundamental nature of what CoT is.

DataAlchemy experiments train LLMs from scratch in controlled environments and probe them under three distributional shift dimensions:

  1. Task distribution shift — novel tasks with unique elements or underlying logical structure not seen during training
  2. Length distribution shift — reasoning chains substantially longer or shorter than training data length range
  3. Format distribution shift — prompt formulation variations (even minor syntactic changes) that fall outside training distribution

In all three dimensions, the pattern is the same: CoT works within distribution, fails outside it. Under moderate shifts, models generate fluent yet logically inconsistent reasoning — the form holds, the logic breaks. This is the "mirage" phenomenon: outputs look like reasoning while producing wrong conclusions.

The interpretive frame: CoT reflects a structured inductive bias learned from training data, not a generalizable reasoning capability. When a test query is within this inductive bias, CoT activates the appropriate reasoning schema and produces good outputs. When the query falls outside it, the schema mismatch produces confident-sounding nonsense.

The practical implication for CoT as a plug-and-play solution: it is not. Performance on CoT benchmarks measures in-distribution capability. Extrapolating to novel tasks, unusual prompt formulations, or unusually long/short reasoning chains is unjustified. The benchmark scores do not predict performance under distribution shift.

This provides the empirical grounding for Does chain-of-thought reasoning reveal genuine inference or pattern matching? — the mirage emerges from imitation under distribution shift: the model continues imitating the form of reasoning while having no schema to produce valid content.

Inquiring lines that read this note 265

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Is language model reasoning authentic and what causes models to reason? How should conversational recommenders balance preference elicitation with direct recommendation? How do capability benchmark scores systematically misrepresent true model abilities? What causes reasoning models to fail or wander off track? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? What is the relationship between thinking tokens and reasoning accuracy? Can models improve accuracy without degrading reasoning quality? Do reasoning benchmarks predict model performance in long-horizon workflows? How does reasoning length affect model performance across different tasks? What reasoning architectures enable models to solve complex problems efficiently? How should systems decide whether to retrieve or reason alone? Can diffusion models match autoregressive performance on language generation tasks? Is reasoning capability latent in base models or created by post-training? How much do training data properties shape model reasoning? Can mechanistic interpretability reliably guide practical model design choices? Can inference-time compute effectively substitute for model scale? How much does training format versus domain influence reasoning? How does the generation-verification gap limit what we can measure about AI reasoning? Why doesn't reasoning volume improve theory of mind performance? Why do some clarifying approaches produce understanding while others just satisfy? Why do stronger reasoning capabilities create tradeoffs with instruction following? What training dynamics and scale trigger emergence of reasoning capabilities? How do neural networks achieve compositional generalization at scale? How should inference compute be allocated based on problem difficulty? What role does sparsity play in model behavior and scaling decisions? How do prompting refinements mask underlying biases and model frequency patterns? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Can reasoning scale in latent space without tokens? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? Do language models reason through causal mechanisms or semantic associations? Do reasoning traces faithfully reflect actual model reasoning? How effectively can language models perform reasoning, especially combined with symbolic methods? Can parallel reasoning outperform sequential reasoning under fixed token budgets? Do language models develop actual world models or merely task heuristics? How do evaluation practices shape which failures stay visible? How do prompt design choices influence model reasoning and performance? What makes distillation transfer some model capabilities while suppressing others? How does policy entropy collapse constrain scaling of reasoning-focused RL? What makes step-level supervision effective for complex reasoning traces? Does RL create genuinely new reasoning capabilities or refine existing ones? How should test-time compute scaling work in agentic systems? What articulatory and acoustic information does speech preserve that transcription destroys? Can multi-agent systems avoid converging on false agreement without deliberation? What enables genuine semantic understanding in language models? Do structural constraints outperform deep architectures in recommendation systems? How do agent-learned skills transfer and improve across different tasks? What makes imperfect LLM judges safe for optimization? How do spurious versus genuine rewards shape model reasoning and behavior? What structural distinctions matter in reasoning and argumentation? What capability trade-offs arise from domain specialization through fine-tuning? Why do token-level mechanisms matter for learning to reason? How does decomposing tasks improve reasoning and prevent failure propagation? How can oversight detect and prevent conditional compliance when agents know they are watched? Does abstract user knowledge outperform concrete interaction history in personalization? Does alignment training create genuine alignment or just output compliance? Can brute-force automated research substitute for iterative depth and human research intuition? What attack surfaces do reasoning traces and chains introduce? Can we reliably detect when models game evaluations? Why does polished presentation create unearned authority in AI outputs? When should work require human-AI partnership versus full automation? What training data selection strategies maximize generalization across difficulty levels? Can memory architectures handle ultra-long context better than attention?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 139 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

cot reasoning is distribution-bounded — effectiveness degrades predictably with distributional discrepancy