SYNTHESIS NOTE
Topics›Flaws›this note

Do language models fail at reasoning due to complexity or novelty?

Explores whether reasoning-model failures stem from task complexity thresholds or from encountering unfamiliar instances. Tests whether scaling chain length actually addresses the root cause of reasoning breakdown.

Synthesis note · 2026-04-07 · sourced from Flaws

The standard narrative around reasoning-model failures — from Shojaee et al.'s Illusion of Thinking onward — frames the phenomenon as a "complexity threshold" or "step threshold": models handle short reasoning chains but break on long ones. Something about the quantity of reasoning breaks down past some limit. The Chollet-Kambhampati exchange reframes this at the instance level, and the reframing matters for what "improving reasoning" can mean.

Chollet's claim: "Many people assume that LRM reasoning breaks down past a certain 'complexity' or 'number of steps' threshold. This is incorrect. It breaks down past an unfamiliarity threshold. And that threshold is very low. There is no limit to the complexity of tasks you can solve with these models, no limit to the number of steps in the reasoning chains they can master — as long as they have been covered during training/tuning. However, show them something unfamiliar, even very simple and requiring just a handful of reasoning steps (e.g., an ARC 2 task), and they will fail." The apparent complexity threshold in Tower of Hanoi exists because Tower of Hanoi is a familiar problem — the step count at which models fail corresponds to the step count at which instances stop appearing in their training data. Scaling step count is an indirect way of generating novelty, not an independent difficulty axis.

Kambhampati adds the systematic observation: LRMs lose accuracy as familiar-problem instances grow because they don't learn algorithms — they fit instance-based patterns. The two agree on the substantive claim even while they initially disagreed on terminology: "We don't actually disagree, we all know that Transformers don't fit generalizable algorithms, they fit instance-based patterns. It doesn't change the fact that the crux of the problem is familiar vs unfamiliar (at the instance level, not at the abstract 'task' level)."

The reframing has sharp implications. First, the intuition that "just scale more reasoning tokens" as a solution to reasoning failures is structurally misguided. If reasoning failure is instance-novelty-driven, then scaling tokens — which extends the reasoning chain — helps only if the longer chain covers more familiar instance territory. It does not extend to any genuinely unfamiliar instance, no matter how short. Second, the natural evaluation target shifts. Benchmarks that scale complexity (Tower of Hanoi with larger N, River Crossing with more pairs) are generating instance novelty indirectly through size. ARC 2 and similar benchmarks generate instance novelty directly through task structure change. The latter is a better measure of whether the model is fitting algorithms or fitting patterns. Third, the definition of "familiarity" matters and Chollet makes it precise: "outside of the classroom, in the real world, you are never exposed to neatly defined 'tasks' and step-by-step algorithms, you are only exposed to situations. Intelligence is the ability to infer generalizable algorithms from situations (instances) only. So the only reasonable definition of familiarity/novelty is at the situation/instance level. If you define it with respect to algorithms you are assuming the problem has already been solved."

This aligns with and sharpens several existing notes. Do foundation models learn world models or task-specific shortcuts? identified task-specific heuristics as the mechanism; Chollet-Kambhampati identify the corresponding failure condition — the heuristics work where they have instance coverage and fail where they do not. Do transformers actually learn systematic compositional reasoning? provides the mathematical substrate: if compositional reasoning is subgraph matching, then novelty at the subgraph level is what breaks the mechanism. Does chain-of-thought reasoning reveal genuine inference or pattern matching? extends this to the performance-vs-reasoning gap: CoT imitates the form of abstract reasoning without performing it, which is exactly why it handles familiar problems at scale but fails on unfamiliar problems at low complexity.

The reframing also creates a tension with some optimistic RL results. Can reinforcement learning discover reasoning strategies base models cannot? shows that extended RL can produce strategies not present in the base model. If reasoning is purely instance-pattern-fitting, where does the novelty in ProRL come from? A reconciliation: RL-discovered "novel strategies" may still be instance-family novelty — the model learns to combine previously separate instance patterns in new ways, producing what looks like strategy but is still pattern composition. This would be genuine progress within the instance-pattern regime without escaping it. A test: take a ProRL-extended model and evaluate it on ARC 2. If the instance-novelty thesis is right, ProRL gains should not transfer to instance-level novelty challenges.

The practical implication for evaluation design is straightforward. Current benchmarks that scale complexity to induce failure are indirectly measuring instance coverage in training data. Benchmarks that induce instance novelty at fixed short complexity — ARC 2, held-out reasoning tasks with genuinely new structure — measure what matters: whether the model is doing anything other than pattern lookup.

Inquiring lines that read this note 262

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What causes reasoning models to fail or wander off track? Why does adding new knowledge through fine-tuning degrade existing capabilities? Why do stronger reasoning capabilities create tradeoffs with instruction following? What is the relationship between thinking tokens and reasoning accuracy? What compositional reasoning failures limit large language models despite scale? Is reasoning capability latent in base models or created by post-training? Do reasoning traces faithfully reflect actual model reasoning? Can reasoning scale in latent space without tokens? What types of diversity prevent reasoning systems from collapsing? How should designers communicate what AI systems truly are and can do? How does self-revision in reasoning models affect accuracy and confidence? How does policy entropy collapse constrain scaling of reasoning-focused RL? How do multi-agent LLM systems fail distinctly compared to single agents? How effectively can language models perform reasoning, especially combined with symbolic methods? Can models improve accuracy without degrading reasoning quality? Do language models learn genuine understanding or just surface patterns? Do language models reason like humans or mimic surface patterns? Is language model reasoning authentic and what causes models to reason? What training data selection strategies maximize generalization across difficulty levels? Why doesn't reasoning volume improve theory of mind performance? Can prompt-based context override biases that were embedded during pretraining? Do reasoning benchmarks predict model performance in long-horizon workflows? What training dynamics and scale trigger emergence of reasoning capabilities? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? How do neural networks achieve compositional generalization at scale? How should inference compute be allocated based on problem difficulty? How do prompting refinements mask underlying biases and model frequency patterns? What causes retrieval-augmented generation systems to fail despite access to external knowledge? What reasoning architectures enable models to solve complex problems efficiently? What enables genuine semantic understanding in language models? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? How much do training data properties shape model reasoning? Does encoded knowledge in language models actually influence their outputs? How does reasoning length affect model performance across different tasks? How do prompt design choices influence model reasoning and performance? What attack surfaces do reasoning traces and chains introduce? Can multi-agent systems avoid converging on false agreement without deliberation? Do language models develop actual world models or merely task heuristics? Do language models reason through causal mechanisms or semantic associations? How do evaluation practices shape which failures stay visible? Why don't LLMs reliably translate capability into accurate outputs? What makes distillation transfer some model capabilities while suppressing others? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How does improved reasoning affect models' ability to acknowledge uncertainty? Can inference-time compute effectively substitute for model scale? Can mechanistic interpretability reliably guide practical model design choices? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Does model confidence reliably signal actual accuracy in practice? What makes step-level supervision effective for complex reasoning traces? What articulatory and acoustic information does speech preserve that transcription destroys? Why do agents falsely report success on failed tasks? How does the generation-verification gap limit what we can measure about AI reasoning? How do surface patterns enable correct outputs but reduce robustness? What structural distinctions matter in reasoning and argumentation? What makes imperfect LLM judges safe for optimization? How does evaluation scope and dimensionality affect what we measure? Why do token-level mechanisms matter for learning to reason? How do capability benchmark scores systematically misrepresent true model abilities? What role does sparsity play in model behavior and scaling decisions? How should retrieval systems handle complex multi-step reasoning? Why do some clarifying approaches produce understanding while others just satisfy? How does synthetic data quality and diversity affect downstream model capabilities? How does decomposing tasks improve reasoning and prevent failure propagation? Can intelligent routing over smaller models outperform scaling a single large model? What mechanisms preserve shared understanding in evolving conversations?

Related concepts in this collection 11

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
24 direct connections · 238 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LRM reasoning breakdown is driven by instance-level unfamiliarity not task-level complexity — there is no limit to reasoning chain length as long as the instances were covered during training