Do language models fail at reasoning due to complexity or novelty?
Explores whether reasoning-model failures stem from task complexity thresholds or from encountering unfamiliar instances. Tests whether scaling chain length actually addresses the root cause of reasoning breakdown.
The standard narrative around reasoning-model failures — from Shojaee et al.'s Illusion of Thinking onward — frames the phenomenon as a "complexity threshold" or "step threshold": models handle short reasoning chains but break on long ones. Something about the quantity of reasoning breaks down past some limit. The Chollet-Kambhampati exchange reframes this at the instance level, and the reframing matters for what "improving reasoning" can mean.
Chollet's claim: "Many people assume that LRM reasoning breaks down past a certain 'complexity' or 'number of steps' threshold. This is incorrect. It breaks down past an unfamiliarity threshold. And that threshold is very low. There is no limit to the complexity of tasks you can solve with these models, no limit to the number of steps in the reasoning chains they can master — as long as they have been covered during training/tuning. However, show them something unfamiliar, even very simple and requiring just a handful of reasoning steps (e.g., an ARC 2 task), and they will fail." The apparent complexity threshold in Tower of Hanoi exists because Tower of Hanoi is a familiar problem — the step count at which models fail corresponds to the step count at which instances stop appearing in their training data. Scaling step count is an indirect way of generating novelty, not an independent difficulty axis.
Kambhampati adds the systematic observation: LRMs lose accuracy as familiar-problem instances grow because they don't learn algorithms — they fit instance-based patterns. The two agree on the substantive claim even while they initially disagreed on terminology: "We don't actually disagree, we all know that Transformers don't fit generalizable algorithms, they fit instance-based patterns. It doesn't change the fact that the crux of the problem is familiar vs unfamiliar (at the instance level, not at the abstract 'task' level)."
The reframing has sharp implications. First, the intuition that "just scale more reasoning tokens" as a solution to reasoning failures is structurally misguided. If reasoning failure is instance-novelty-driven, then scaling tokens — which extends the reasoning chain — helps only if the longer chain covers more familiar instance territory. It does not extend to any genuinely unfamiliar instance, no matter how short. Second, the natural evaluation target shifts. Benchmarks that scale complexity (Tower of Hanoi with larger N, River Crossing with more pairs) are generating instance novelty indirectly through size. ARC 2 and similar benchmarks generate instance novelty directly through task structure change. The latter is a better measure of whether the model is fitting algorithms or fitting patterns. Third, the definition of "familiarity" matters and Chollet makes it precise: "outside of the classroom, in the real world, you are never exposed to neatly defined 'tasks' and step-by-step algorithms, you are only exposed to situations. Intelligence is the ability to infer generalizable algorithms from situations (instances) only. So the only reasonable definition of familiarity/novelty is at the situation/instance level. If you define it with respect to algorithms you are assuming the problem has already been solved."
This aligns with and sharpens several existing notes. Do foundation models learn world models or task-specific shortcuts? identified task-specific heuristics as the mechanism; Chollet-Kambhampati identify the corresponding failure condition — the heuristics work where they have instance coverage and fail where they do not. Do transformers actually learn systematic compositional reasoning? provides the mathematical substrate: if compositional reasoning is subgraph matching, then novelty at the subgraph level is what breaks the mechanism. Does chain-of-thought reasoning reveal genuine inference or pattern matching? extends this to the performance-vs-reasoning gap: CoT imitates the form of abstract reasoning without performing it, which is exactly why it handles familiar problems at scale but fails on unfamiliar problems at low complexity.
The reframing also creates a tension with some optimistic RL results. Can reinforcement learning discover reasoning strategies base models cannot? shows that extended RL can produce strategies not present in the base model. If reasoning is purely instance-pattern-fitting, where does the novelty in ProRL come from? A reconciliation: RL-discovered "novel strategies" may still be instance-family novelty — the model learns to combine previously separate instance patterns in new ways, producing what looks like strategy but is still pattern composition. This would be genuine progress within the instance-pattern regime without escaping it. A test: take a ProRL-extended model and evaluate it on ARC 2. If the instance-novelty thesis is right, ProRL gains should not transfer to instance-level novelty challenges.
The practical implication for evaluation design is straightforward. Current benchmarks that scale complexity to induce failure are indirectly measuring instance coverage in training data. Benchmarks that induce instance novelty at fixed short complexity — ARC 2, held-out reasoning tasks with genuinely new structure — measure what matters: whether the model is doing anything other than pattern lookup.
Inquiring lines that read this note 262
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What causes reasoning models to fail or wander off track?- How can minimal pairs expose reasoning failures that single-instance accuracy metrics miss?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- How does the frame problem differ between symbolic and statistical reasoning systems?
- Why do reasoning models fail on structurally unfamiliar instances?
- Does text-only evaluation hide reasoning collapse that tool use could repair?
- How do humans and LMs differ on multi-hop reasoning?
- Where do humans and language models actually diverge in reasoning ability?
- What makes Compound-QA expose weaknesses in monologue reasoning?
- Why do models automatically adjust reasoning length to problem difficulty?
- How do search tasks differ from derivation tasks in reasoning efficiency?
- What causes snowball errors to accumulate across reasoning steps in language models?
- Why does comparison reasoning generalize better than composition reasoning?
- How does reasoning instability prevent models from modeling individuals?
- Why do reasoning models fail when input length increases even below context limits?
- Why do reasoning chains degenerate into undirected exploration at scale?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- What explains the gap between perplexity performance and actual reasoning capability?
- Can scaffolding frameworks isolate inductive reasoning from deductive confounds?
- Why do reasoning models wander instead of searching systematically?
- Why does cross-text analogical reasoning fail when semantics decouple from symbols?
- Why does removing semantic content collapse reasoning in language models?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- Do reasoning failures stem from strategy or from calculation breakdown?
- Do reasoning models switch approaches when encountering local difficulty?
- What mechanisms cause reasoning models to wander rather than focus?
- How do single wrong steps corrupt entire reasoning chains?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- What failure modes emerge when scheme classification feeds downstream reasoning pipelines?
- Why do language models struggle with backward reasoning compared to forward?
- Can cognitive scaffolding replace tool-based reasoning augmentation in language models?
- What causes reasoning quality to degrade during long research tasks?
- Why do smaller models lose reasoning faithfulness more than larger models?
- Why does second-hop reasoning fail when composed with out-of-distribution triples?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- What makes deterministic recursive reasoning models underperform on multi-solution tasks?
- Why does single-shot learning fail in REVTHINK's multi-source reasoning tasks?
- Is reasoning failure caused by task complexity or training distribution gaps?
- Why does strategy diversity within reasoning chains improve model generalization?
- How does recombining partial trajectories maintain coherence in natural language reasoning?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- How does instance novelty rather than chain length explain reasoning failure?
- What distinguishes the convergence patterns between reasoning and lexical variation tasks?
- Why do structured reasoning representations sometimes reduce rather than improve error detection?
- Can weak models reason better when freed from cognitive load by structure?
- When does knowledge activation fail across different model architectures?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- How does the knowing-doing gap widen as tasks become more complex?
- Why do models fail on logically equivalent tasks with different data distributions?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Is the reasoning cliff actually a tool-use problem?
- Why do difficult problems force models to develop reasoning strategies?
- Can reasoning models succeed at logic but fail at execution?
- Why do reasoning model failures stem from execution rather than reasoning?
- Can machine learning encode pragmatic reasoning about when rules should bend?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- How do reasoning-related features behave when trained on near-impossible problems?
- Can models distinguish between logical impossibility and their own execution limits?
- Why does target probability matter more than task logical complexity?
- How can we turn reasoning model failures into useful training signals?
- Do models genuinely reason harder on difficult tasks or just appear to?
- How does question difficulty and breadth affect what models learn to reason?
- How does learnability at the observer's current state prevent novelty from breaking model reasoning?
- Why does scaling reasoning tokens fail to improve unfamiliar tasks?
- Why do models overthink easy problems and underthink difficult ones?
- How much does schema bloat actually degrade reasoning in large language models?
- Does task difficulty alone determine how many thinking tokens a model should use?
- What limits external scaling when a model lacks reasoning foundation?
- Why do harder puzzles cause all models to collapse despite larger token budgets?
- What makes a problem instance unfamiliar to a language model?
- Does scaling model size solve compositional generalization problems?
- Why do language models fail at planning despite understanding strategies?
- Why do language models fail when semantic content is stripped away?
- Can simple diagnostic tests predict language model performance in production complexity?
- How do rare linguistic registers differ from conceptually complex examples?
- Why do large language models fail at temporal reasoning in complex legal cases?
- What happens when formal languages satisfy hierarchy but fail learnability constraints?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- Why do large language models still have systematic blind spots with complex structures?
- Why do standard NLP benchmarks hide the most critical language limitations?
- Why do language models struggle with formal logical reasoning and joins?
- Why do language models plateau at 55 to 60 percent constraint satisfaction?
- Why do language models fail at understanding ambiguous or complex requirements?
- Why do long-context language models struggle with compositional reasoning tasks?
- Why do language models plateau at constraint satisfaction regardless of scale?
- Are newer larger language models actually worse at faithful summarization?
- What makes domain-specific utterance resolution harder for general large models?
- Why do single examples trigger large reasoning improvements in models?
- Can models learn to select exemplars based on reasoning skills rather than complexity?
- What makes reasoning-specific post-training different from standard parameter scaling?
- How does a single training example trigger phase transitions in reasoning output?
- How do single training examples activate reasoning capabilities in language models?
- Do base models truly possess latent reasoning capability?
- What does pass@k reveal about base model reasoning capacity?
- Can models possess latent reasoning capability that training signals fail to unlock?
- Why do simple length heuristics outperform sophisticated semantic methods?
- What behavioral markers signal when reasoning chains are performative?
- Why do models show performative reasoning on easy tasks but genuine reasoning on hard ones?
- Why do reasoning models produce unfaithful or unhelpful reasoning traces?
- What makes some sentences in reasoning traces have disproportionate causal influence?
- What metric distinguishes deep reasoning from superficial information propagation?
- Why do language models produce reasoning traces that mimic human reasoning style?
- What causes reasoning loops and distraction in models with very long traces?
- Can latent reasoning architectures work as retrofits to existing models?
- Does the latent-explicit gap widen beyond 3B parameters on reasoning tasks?
- Does optimizing directly for semantic diversity improve both reasoning quality and exploration?
- How does majority voting fail when reasoning samples lack genuine diversity?
- Do rare cultural concepts fail predictably as model scale increases?
- Does self-revision actually improve reasoning in large language models?
- Why do models confabulate inconsistently across different samples?
- How does policy entropy collapse constrain token-level distribution in reasoning?
- Why does policy entropy collapse limit reasoning and dialogue RL scaling?
- Can symbolic solvers rescue language models from logical reasoning failures?
- Can reasoning chains work without logical validity?
- How does context complexity affect LLM performance on temporal reasoning tasks?
- How does semantic reasoning differ from symbolic reasoning in language models?
- What makes deductive reasoning so brittle in language models overall?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- How do game-based benchmarks reveal reasoning fragmentation across domains?
- Why do format and structure matter more than actual content in reasoning?
- Why does premise ordering shift syllogistic reasoning performance by over 30 percent?
- Can language models perform purely symbolic reasoning when semantics are removed?
- Can reasoning benchmarks separate logic from believability?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Does model scaling improve knowledge storage faster than reasoning ability?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- How does fine-tuning on natural language inference affect fallacy susceptibility?
- Can reasoning models distinguish between new evidence and manipulative reframing?
- Can dataset design systematically expand reasoning graph diameter?
- Why does naive randomness fail to improve stochastic latent reasoning models?
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- Can expert-derived knowledge bases scale to other high-stakes domains?
- How does evaluation setting affect measured reasoning capabilities in language models?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Why does compositional reasoning fail to explain cross-domain transfer?
- Can language models reason without relying on learned semantic patterns?
- Why do language models fall back on frequency heuristics under structural complexity?
- Why do language models imitate reasoning form without abstract inference capability?
- Do language models learn surface patterns that appear generalizable but actually fail under shift?
- Do language models build world models or just task-specific heuristics?
- Is confabulation inevitable in large language models regardless of training?
- Do language models systematically overestimate accuracy on collective behavior tasks?
- Can benchmark performance distinguish surface from structural linguistic knowledge?
- Why do surface generalizations fail on unusual syntactic structures?
- Can language models reason without relying on surface level pattern matching?
- Is gradient behavior in language functional or a sign of ambiguity?
- Does directional knowledge failure indicate shallow pattern matching over deep representation?
- What causes language models' strategic rationality to decline with increased game complexity?
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- Why does exemplar performance vary across order complexity diversity and style?
- Why do reasoning models perform poorly at theory of mind tasks?
- Why do reasoning models perform worse on theory of mind tasks?
- What makes reasoning models worse at understanding people?
- Why does homework adherence remain low despite advances in language model capability?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Can a complexity-predictor be meaningful if models are redundant?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- What role does curriculum design play in reasoning emergence?
- How sensitive is analogical reasoning emergence to training data and scale?
- Does more inference compute help reasoning models match specialized domain performance?
- How should inference budget adapt based on problem difficulty?
- How should inference budgets adapt based on prompt difficulty?
- Can weaker models match stronger ones with sufficient search and reasoning budget?
- Does irrelevant context degrade reasoning even within model context limits?
- Why do semantically related prompts converge into attractor states in middle layers?
- Does architectural design matter more than model scale for reasoning tasks?
- Can small models solve complex tasks using externalized reasoning graphs?
- Do reasoning systems reuse cognitive structures across unrelated topics?
- Does small-world structure in reasoning graphs improve generalization?
- What computational structures can actually scale serial reasoning depth?
- Can frame semantics explain why context matters more than word similarity?
- What distinguishes conceptual understanding from statistical pattern matching in models?
- Why can't pattern-matching systems perform the observation that expert communication requires?
- Why do multimodal models fail on rare and underrepresented concepts?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- Why do non-experts default to familiar chart types despite domain complexity?
- Why do language models fail at grounding and inference?
- What reveals the epistemic limits of language models?
- Can language models distinguish between novel insight and unjustified conceptual blending?
- Can language models accurately evaluate the quality of their own reasoning?
- What sparse mechanistic structures drive reasoning traces in language models?
- What geometric structure do language models actually use during inference?
- How do semantic and symbolic reasoning capabilities differ in language models?
- How does tool-based reasoning expand what language models can do?
- Why does answer-confirmation bias emerge in language model reasoning?
- Why do simple math problems get worse with longer reasoning chains?
- Does more thinking always help large language models or sometimes hurt?
- How does random walk length control reasoning complexity in question generation?
- Why do longer reasoning chains signal hesitation rather than depth?
- Does distillation from reasoning models spread overthinking to smaller models?
- Why does chain-of-thought prompting fail to fix length-induced reasoning degradation?
- How do longer reasoning chains create vulnerability to attacks?
- Can external classifiers reliably decide when a model should reason?
- Why do longer reasoning chains correlate with lower accuracy in o1-like models?
- Can memorization scores diagnose where reasoning chains become unreliable?
- Why do longer reasoning chains explore like tourists instead of scientists?
- Why do thinking models execute longer tasks than standard language models?
- Why do language models overthink simple questions when given extra time?
- Can removing hierarchy from dual-recurrence models improve reasoning performance?
- Does adding reasoning to models degrade other capabilities like rule inference?
- Why does reasoning performance degrade as input length increases?
- Does longer reasoning always improve model accuracy on complex tasks?
- How should reasoning prompts adapt based on question complexity and type?
- How do exemplar properties affect the brittleness of chain-of-thought prompting?
- How do input length and context size separately affect reasoning quality?
- How do foundation models develop task-specific heuristics instead of world models?
- Why do epistemic failure modes cluster around world model limitations?
- What are collider structures and why do they reveal reasoning errors?
- Why do causal reasoning directions succeed while temporal reasoning directions fail?
- Where do collider-type reasoning errors appear in real-world decisions?
- Why do unresolved items cluster in structured patterns rather than randomly?
- Why do different reasoning chains surface different relevant facts?
- Can parallel reasoning chains outperform longer sequential chains with the same compute?
- Are some problems fundamentally unsolvable by parallel inference methods?
- Do models excel at reasoning depth or memory breadth when scaling test time compute?
- Can reasoning models outperform non-reasoning models with more inference compute?
- Why does chain of thought reasoning fail across different prompt formats?
- Why do verbalized reasoning chains fail on certain problem classes?
- Can instance-adaptive reasoning happen without sequential token dependencies?
- How does making implicit reasoning requirements explicit change model performance?
- How brittle are chain-of-thought exemplars across order and complexity?
- Is verbalized chain-of-thought necessary for language model reasoning?
- Does reasoning training create blind spots in premise detection?
- How does model confidence relate to exemplar brittleness in chain-of-thought?
- Does premature confidence signal flawed reasoning in language models?
- How can high benchmark performance mask broken reasoning in AI systems?
- Does epistemic narrowness appear equally across professions, proofs, and other reasoning tasks?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How do frontier models maintain agreement scores above 90 percent across reasoning tasks?
Related concepts in this collection 11
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
the mechanism beneath the phenomenon; heuristics work within instance coverage and fail outside
-
Do transformers actually learn systematic compositional reasoning?
Explores whether transformers solve compositional tasks through genuine systematic reasoning or by pattern-matching against training data. This matters because it determines whether scaling alone can achieve robust generalization.
the mathematical substrate: subgraph matching is instance-level pattern matching
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
CoT imitates form without performing inference; unfamiliarity reveals the imitation
-
Does more thinking time always improve reasoning accuracy?
Explores whether extending a model's thinking tokens linearly improves performance, or if there's a point beyond which additional reasoning becomes counterproductive.
the apparent threshold may be unfamiliarity not tokens
-
Why do reasoning LLMs fail at deeper problem solving?
Explores whether current reasoning models systematically search solution spaces or merely wander through them, and how this affects their ability to solve increasingly complex problems.
wandering may be the novelty response
-
Does the reasoning cliff depend on how we test models?
If language models hit a capability wall in text-only reasoning tasks, does that wall disappear when they can use tools? What does this reveal about what we're actually measuring?
complementary reframing at the execution layer
-
Can reinforcement learning discover reasoning strategies base models cannot?
Does RL training truly expand what models can do, or does it just find solutions already hidden in base models? ProRL tests this by running RL longer and on diverse tasks beyond mathematics.
apparent tension; possibly resolved as instance-family novelty rather than algorithm novelty
-
Can neural networks learn compositional skills without symbolic mechanisms?
Do neural networks need explicit symbolic architecture to compose learned concepts, or can scaling alone enable compositional generalization? This asks whether compositionality is an architectural feature or an emergent property of scale.
partial counterpoint: scaling data closes some generalization gaps, but instance novelty remains the boundary
-
Can identical outputs hide broken internal representations?
Can neural networks produce correct outputs while having fundamentally fractured internal structure that prevents generalization and creativity? This challenges our assumptions about what performance benchmarks actually measure.
FER is the representation-level parallel; identical benchmark scores can mask different instance coverage
-
Can transformers improve exponentially by learning from their own correct solutions?
Can standard transformers achieve extreme length generalization by iteratively filtering and training on their own correct outputs? This explores whether self-correction loops enable unbounded out-of-distribution improvement without architectural changes.
subtle counterpoint: length generalization within a familiar task family (addition at longer digit counts) still extends beyond initial instance coverage through iteration; but the instance type stays familiar, so this may be "same-algorithm novelty" that the thesis accommodates
-
Are reasoning model collapses really failures of reasoning?
Explores whether language models hit a fundamental reasoning ceiling or whether text-only evaluation masks execution limitations. Examines how tool access might reveal hidden reasoning capabilities.
alternative diagnosis at the execution layer
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Large Language Model Reasoning Failures
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
Original note title
LRM reasoning breakdown is driven by instance-level unfamiliarity not task-level complexity — there is no limit to reasoning chain length as long as the instances were covered during training