Does chain-of-thought reasoning actually generalize beyond training data?
Explores whether CoT's strong performance on benchmarks reflects genuine reasoning ability or merely reflects learned patterns tied to specific distributions. Tests how CoT behaves when tasks, formats, or reasoning length shift away from training data.
Chain-of-Thought prompting performs well on in-distribution problems and fails predictably as distributional discrepancy increases. This is not a bug — it is the fundamental nature of what CoT is.
DataAlchemy experiments train LLMs from scratch in controlled environments and probe them under three distributional shift dimensions:
- Task distribution shift — novel tasks with unique elements or underlying logical structure not seen during training
- Length distribution shift — reasoning chains substantially longer or shorter than training data length range
- Format distribution shift — prompt formulation variations (even minor syntactic changes) that fall outside training distribution
In all three dimensions, the pattern is the same: CoT works within distribution, fails outside it. Under moderate shifts, models generate fluent yet logically inconsistent reasoning — the form holds, the logic breaks. This is the "mirage" phenomenon: outputs look like reasoning while producing wrong conclusions.
The interpretive frame: CoT reflects a structured inductive bias learned from training data, not a generalizable reasoning capability. When a test query is within this inductive bias, CoT activates the appropriate reasoning schema and produces good outputs. When the query falls outside it, the schema mismatch produces confident-sounding nonsense.
The practical implication for CoT as a plug-and-play solution: it is not. Performance on CoT benchmarks measures in-distribution capability. Extrapolating to novel tasks, unusual prompt formulations, or unusually long/short reasoning chains is unjustified. The benchmark scores do not predict performance under distribution shift.
This provides the empirical grounding for Does chain-of-thought reasoning reveal genuine inference or pattern matching? — the mirage emerges from imitation under distribution shift: the model continues imitating the form of reasoning while having no schema to produce valid content.
Inquiring lines that read this note 265
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Is language model reasoning authentic and what causes models to reason?- What makes conceptual inquiry the fastest high-scoring AI interaction pattern?
- Can chain of thought reasoning actually validate logical arguments?
- What other hidden biases might aggregate metrics fail to distinguish from reasoning?
- How should we redesign benchmarks to catch conservative bias in reasoning tasks?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How do frontier models maintain agreement scores above 90 percent across reasoning tasks?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- How can hidden test partitions detect constant predictions that generalize?
- Can surface heuristics override implicit constraints in domain-specific reasoning?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- What makes a background condition relevant to a specific reasoning task?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- How do humans and LMs differ on multi-hop reasoning?
- Why does comparison reasoning generalize better than composition reasoning?
- Why are pairwise relations insufficient for representing higher-order multi-hop reasoning?
- What explains the gap between perplexity performance and actual reasoning capability?
- Can scaffolding frameworks isolate inductive reasoning from deductive confounds?
- Does this reasoning steering method work consistently across all model sizes?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Why does second-hop reasoning fail when composed with out-of-distribution triples?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- Why does single-shot learning fail in REVTHINK's multi-source reasoning tasks?
- Is reasoning failure caused by task complexity or training distribution gaps?
- Why does strategy diversity within reasoning chains improve model generalization?
- How does instance novelty rather than chain length explain reasoning failure?
- What distinguishes the convergence patterns between reasoning and lexical variation tasks?
- Does chain-of-thought text causally drive reasoning or merely reflect it?
- Can steering a single latent feature replicate chain-of-thought performance?
- What detection methods can catch each distinct CoT bypass strategy?
- Does changing decoding procedure reveal hidden chain-of-thought paths?
- Why does chain-of-thought fail when problems lack matching training schemata?
- Is chain-of-thought reasoning actual computation or distribution imitation?
- What happens to chain-of-thought performance across distribution shifts?
- How does chain-of-thought training change higher layer computations?
- Does chain-of-thought reasoning specifically improve performance on metalinguistic tasks?
- Does chain-of-thought reasoning improve mental state tracking in dialogue?
- Do chain-of-thought explanations reveal genuine reasoning or trigger latent features?
- Why do we measure reasoning quality by reading visible chains?
- Why does long CoT training optimize for structural coherence over content correctness?
- How does chain-of-thought reasoning become decorative after domain-specific fine-tuning?
- Does chain-of-thought reasoning help or hurt social reasoning tasks?
- Why does chain-of-thought fail to improve multimodal model perception performance?
- Does CoT reasoning actually cause the outputs that follow it?
- Can single representation edits match chain-of-thought reasoning without explicit steps?
- Does reasoning training create blind spots in premise detection?
- How does trajectory geometry relate to the need for chain-of-thought reasoning?
- Why does scaling reasoning tokens fail to improve unfamiliar tasks?
- How much does test-time compute improve reasoning without more tokens?
- Can activation steering vectors compress reasoning without retraining models?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- Can reasoning benchmarks separate logic from believability?
- Can activation patching reveal which reasoning steps actually matter?
- How much does pre-training frequency predict reasoning task performance?
- Does domain training degrade reasoning ability even when benchmark scores rise?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- How does cross-domain reasoning transfer differ from domain-specific knowledge transfer?
- Does model scaling improve knowledge storage faster than reasoning ability?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Can dataset design systematically expand reasoning graph diameter?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Can attribute decomposition improve other interactive reasoning tasks beyond clinical questioning?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- Can benchmark improvements hide degradation of deliberative reasoning?
- Can mathematical reasoning improvements transfer across problem subdomains?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- What makes procedural knowledge in documents generalize better than facts?
- Can expert-derived knowledge bases scale to other high-stakes domains?
- Can small demonstration sets unlock general reasoning without large question data?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Why do reasoning gains resist clear attribution to specific training changes?
- Does the Heuristic Override Benchmark measure enumeration or world knowledge?
- Do synthetic verification chains from long-CoT models match the quality of human-annotated process labels?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Can step-level deliberation flags guide other reasoning systems?
- Do explicit reasoning chains improve or harm performance on complex judgment tasks?
- Can extended thinking genuinely improve reasoning or just increase variance?
- How do gradient descent iterations at inference compare to chain-of-thought reasoning chains?
- When does explicit reasoning actually degrade performance on a task?
- How do chain-of-thought structures affect reasoning robustness?
- Can extended reasoning training capture individual strategic thinking styles?
- Why does extended thinking increase output variance without improving reasoning quality?
- Does deep-thinking ratio measure computational effort better than chain-of-thought length?
- Can models trained on longer contexts develop better fundamental reasoning abilities?
- Why do longer reasoning chains correlate with lower accuracy in o1-like models?
- How much reasoning depth do we actually need for most real-world tasks?
- Does penalizing thought transitions improve reasoning without model retraining?
- Can we improve reasoning by amplifying information at mutual information peaks?
- Can memorization scores diagnose where reasoning chains become unreliable?
- Can minimal reasoning steps match verbose reasoning accuracy?
- Why does extended chain-of-thought reasoning fail to improve numerical optimization performance?
- Does chain-of-thought accuracy degrade with longer reasoning traces?
- Can graph cyclicity and topology predict when reasoning systems achieve breakthrough insights?
- What formal representation could capture analogical reasoning across domains?
- Can small models solve complex tasks using externalized reasoning graphs?
- Can we transfer reasoning structure without copying surface form?
- Do reasoning systems reuse cognitive structures across unrelated topics?
- Can recursive subtask trees implement tree-of-thought reasoning more efficiently?
- How does graph of thoughts enable divide-and-conquer reasoning patterns?
- What makes multi-paradigm chaining a distinct reasoning topology?
- Does small-world structure in reasoning graphs improve generalization?
- What makes a causal abstraction more transferable than a generic heuristic?
- What computational structures can actually scale serial reasoning depth?
- How does structured environment state compare to transcript replay for multi-turn reasoning?
- What role does embedding space geometry play in multi-hop reasoning?
- How can per-step decisions about knowledge retrieval improve reasoning over uniform policies?
- How do retrieval heads enable chain-of-thought reasoning to reference earlier context?
- What makes reasoning capability a pre-training rather than post-training phenomenon?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Can reasoning skills trained on law improve performance in STEM?
- What makes reasoning-specific post-training different from standard parameter scaling?
- How does a single training example trigger phase transitions in reasoning output?
- How can one training example improve reasoning across thousands of unseen problems?
- What makes thought identifiability provable without auxiliary training data?
- Why does reasoning training improve math but hurt knowledge tasks?
- Does latent reasoning capability exist in base models before any training?
- How do timing and search internalization interact during reasoning post-training?
- Why does reasoning transfer across different numbers but factual recall does not?
- Can smaller amounts of diverse reasoning demonstrations replace exhaustive factual training data?
- What makes token-level reasoning during pretraining different from test-time chain-of-thought?
- Does token-level reasoning during pretraining improve general reasoning without task-specific supervision?
- How does RPT compare to learning when versus how to deploy reasoning?
- Can distillation from stronger models create genuinely new reasoning abilities?
- What does pass@k reveal about base model reasoning capacity?
- What makes some reasoning strategies genuinely novel versus latent?
- How much of Occamy's result comes from training versus the base model?
- Does looped pretraining build reasoning more efficiently than supervised fine-tuning?
- How does cognitive fit theory explain why different tasks need different knowledge structures?
- How much of the combinatorial task space must training data cover?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- How does training data distribution create asymmetric competence across relation types?
- What makes certain bond distributions more learnable than others?
- How much does training composition affect syntactic versus reasoning performance?
- Why does semantic similarity retrieval enable skill transfer to novel situations?
- What distinguishes data that generalizes broadly from task-specific memorization?
- How does training data structure shape reasoning strategy more than domain content?
- Why do non-experts default to familiar chart types despite domain complexity?
- Can scaling data alone solve performance gaps on long-tail concepts?
- Can curriculum graphs as training data improve model understanding of prerequisite chains?
- How does inference compute substitution affect the training parameter scaling trade-off?
- Can test-time scaling prioritize genuine reasoning over pattern matching?
- Does inference-time compute improve pretraining data efficiency in practice?
- What inference-time scaling benefits emerge from reasoning before each prediction?
- What patterns emerge across test-time scaling and reasoning architectures?
- Can reasoning models outperform non-reasoning models with more inference compute?
- How much does training data format shape what reasoning strategy emerges?
- Why does training format shape reasoning strategy more than domain?
- Why does training data format shape reasoning strategy more than domain content?
- Does training data format shape model reasoning more than domain content?
- How does training format shape reasoning strategy more than content?
- How does training data format shape whether models reason in parallel or sequentially?
- How much does training data presentation format shape reasoning ability?
- How does training data format shape which reasoning patterns emerge in models?
- Why does training data format shape reasoning strategy more than content?
- Can training format itself shape what reasoning strategy a model learns?
- Does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- How do surface correlations between narratives and answers mislead benchmark validity?
- How should domain-specific AI be evaluated differently from general benchmarks?
- Why do AI benchmarks measure accuracy instead of reasoning quality?
- How can high benchmark performance mask broken reasoning in AI systems?
- What distribution patterns appear across different theory-of-mind datasets?
- Can theory of mind models generalize across structurally similar scenarios?
- Can structured theory of mind benchmarks measure genuine mental state reasoning?
- Why does fine-tuning sometimes damage chain-of-thought reasoning even when accuracy improves?
- Why do models fail on logically equivalent tasks with different data distributions?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Does reasoning structure match explicit versus implicit task demands?
- Does scaling reasoning capability create tradeoffs with instruction following?
- Is the reasoning cliff actually a tool-use problem?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Why do instruction following and reasoning capability trade off in training?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- How sensitive is analogical reasoning emergence to training data and scale?
- Can adaptive compute distribution across prompts replace the need for sophisticated reasoning frameworks?
- Does more inference compute help reasoning models match specialized domain performance?
- What mechanisms drive test-time compute allocation in reasoning tasks?
- Why do task-specific heuristics fail at generalizing to sparse data regions?
- Does sparsity-guided ordering work equally well for reasoning and classification tasks?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Can continuous latent reasoning match discrete chain-of-thought without training modifications?
- Can articulating latent reasoning processes improve transfer across domains?
- Can latent reasoning scale test-time compute without verbalized tokens or special training?
- Does the latent-explicit gap widen beyond 3B parameters on reasoning tasks?
- Can hyperedges replace triple-based externalization in reasoning tasks?
- Can knowledge graphs externalize and validate reasoning steps during inference?
- Can knowledge graph structure alone generate sufficient training signals for domain reasoning?
- How do random walk reasoning chains from knowledge graphs compare to traditional fine-tuning?
- What saliency patterns distinguish successful from failed chain-of-thought reasoning?
- How does post-training on traces improve performance without semantic reasoning?
- Can deliberate corruption of reasoning traces harm out of distribution generalization?
- What metric distinguishes deep reasoning from superficial information propagation?
- Does trace length actually reflect problem difficulty or training proximity?
- Do longer chain-of-thought traces improve interpretability or just performance?
- How much do compressed reasoning traces transfer across different models?
- Why do shorter confident reasoning traces fail on out-of-distribution problems?
- Can post-hoc analysis of reasoning traces actively mislead users?
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- Can reasoning traces that feel convincing fail to help people predict behavior?
- Do task-specific heuristics emerge because they compress well enough?
- How do game-based benchmarks reveal reasoning fragmentation across domains?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Can breadth-first search in continuous space outperform chain-of-thought on logical tasks?
- How does meta-reasoning combine information distributed across multiple chains?
- How does MCTS combine parallel exploration with sequential reasoning depth?
- Can frozen world models from training cutoff remain adequate for real-world reasoning?
- What distinguishes task-specific heuristics from genuine world models?
- Why does distillation transfer reasoning patterns with few examples?
- Does reasoning style transfer matter more than solution correctness in distillation?
- Does policy entropy collapse limit how many iterations of reasoning training work?
- Does policy entropy collapse in formal reasoning produce the same outcome in social reasoning?
- Does stable entropy in policy training actually guarantee stable reasoning behavior?
- How does reinforcement learning differ from chain-of-thought distillation?
- How does RL compress reasoning path diversity during training?
- Can base models spontaneously produce reasoning traces without any RL training?
- How do extrapolative and contextual generalization measure RL reasoning gains?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
DataAlchemy provides the empirical confirmation: imitation fails under distribution shift because no schema matches
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
distribution-bounded CoT is neither sufficient (fails under shift) nor necessary (in-distribution performance may not require the chain)
-
Can models pass tests while missing the actual grammar?
Do language models succeed on grammatical benchmarks by learning surface patterns rather than structural rules? This matters because correct outputs may hide reliance on shallow heuristics that fail on novel structures.
same pattern: surface patterns work in-distribution, fail under structural change
-
Does training data format shape reasoning strategy more than domain?
What explains why models trained on multiple-choice data reason differently than those trained on free-form text? The research isolates format and domain effects to measure which one matters more.
format-dependency is part of distribution-boundedness: changing the format is a distribution shift
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Hierarchical Reasoning Model
- Break the Chain: Large Language Models Can be Shortcut Reasoners
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Chain of Thoughtlessness? An Analysis of CoT in Planning
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
Original note title
cot reasoning is distribution-bounded — effectiveness degrades predictably with distributional discrepancy