Why do reasoning LLMs fail at deeper problem solving?
Explores whether current reasoning models systematically search solution spaces or merely wander through them, and how this affects their ability to solve increasingly complex problems.
"Reasoning LLMs are Wandering Solution Explorers" provides the most rigorous formalization yet of why reasoning models fail as problem complexity increases. The claim: current RLLMs do not systematically explore solution spaces. They wander.
Systematic exploration requires three properties: (a) validity — the trace follows the reachability structure; (b) effectiveness — the trace contains at least one goal state; (c) necessity — every state in the trace contributes to goal discovery or dead-end elimination. Current models fail all three.
The formalization makes the failure quantifiable. A wandering RLLM performing depth-first search on a binary tree of depth d has a probability pw of omitting one of two child nodes at each decision point. The success probability drops exponentially with depth d. This is not a gradual degradation — it is catastrophic. Problems that appear within reach at depth 5 become virtually impossible at depth 15 not because the model lacks reasoning ability but because it lacks search discipline.
Four failure modes are identified:
- Invalid exploration: transitions violate the problem's reachability structure
- Unnecessary exploration: superfluous states that don't contribute to goal discovery
- Evaluation error: misinterpreting current state or executing planned moves erroneously
- Hallucinated conclusions: claiming solutions that don't satisfy problem constraints
The finding directly challenges the "more thinking tokens = better reasoning" narrative. A wandering model given more tokens doesn't explore more systematically — it wanders more extensively. This is the mechanism behind Does more thinking time always improve reasoning accuracy?: additional compute doesn't fix structural search deficiency.
The exponential degradation result connects to Does policy entropy collapse limit reasoning performance in RL?. Entropy collapse reduces exploration diversity during training; wandering reduces exploration discipline during inference. Both are manifestations of the same problem: the model converges on familiar patterns rather than systematically covering the solution space.
Apple's three-regime confirmation. "The Illusion of Thinking" (Apple) provides independent confirmation through controllable puzzle environments with precise complexity manipulation. Three performance regimes emerge: (1) low-complexity — standard models outperform reasoning models with greater token efficiency; (2) medium-complexity — reasoning models gain advantage through extended thinking; (3) high-complexity — both model types collapse to zero. Near the collapse point, reasoning models reduce their reasoning effort despite having ample token budget — a counterintuitive behavioral scaling limit. Even providing explicit optimal algorithms does not prevent collapse, confirming the bottleneck is execution not conceptualization. The three-regime structure refines the wandering explorer thesis: wandering is harmful at low complexity (overthinking easy problems), partially beneficial at medium complexity (exploring toward solutions), and irrelevant at high complexity (no amount of wandering reaches the goal).
Inquiring lines that read this note 112
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do language models develop actual world models or merely task heuristics? What causes reasoning models to fail or wander off track?- Can surface heuristics override implicit constraints in domain-specific reasoning?
- Why do reasoning models fail on structurally unfamiliar instances?
- How do humans and LMs differ on multi-hop reasoning?
- Why do models automatically adjust reasoning length to problem difficulty?
- How do search tasks differ from derivation tasks in reasoning efficiency?
- Why does comparison reasoning generalize better than composition reasoning?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Do reasoning models overthink ill-posed questions instead of recognizing incompleteness?
- What explains the gap between perplexity performance and actual reasoning capability?
- Why do reasoning models wander instead of searching systematically?
- What distinguishes systematic search from wandering exploration in reasoning?
- Do reasoning failures stem from strategy or from calculation breakdown?
- Do reasoning models switch approaches when encountering local difficulty?
- What mechanisms cause reasoning models to wander rather than focus?
- What failure modes emerge when scheme classification feeds downstream reasoning pipelines?
- What happens to iterative search quality when reasoning depth is unconstrained?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- When is numeric computation the real bottleneck versus reasoning depth?
- What makes deterministic recursive reasoning models underperform on multi-solution tasks?
- How does active reasoning through interaction differ from passive single-turn problem solving?
- Is reasoning failure caused by task complexity or training distribution gaps?
- How does o1-style reasoning relate to learned search processes versus memorized solutions?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- Where do LLMs succeed at generation but struggle with evaluation?
- Why do LLM personas struggle with specificity in specialized domains like law?
- What specific execution barriers do LLM ideas encounter most frequently?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Can LLMs explain concepts correctly while failing to use them?
- Why do LLMs plateau on creativity tasks while humans reach further?
- Where do LLMs fail as knowledge systems compared to humans?
- Why do LLMs generate novel ideas but lack evaluative commitment?
- Do LLMs fail exploration because of context integration or computational limitations?
- Which knowledge types do LLMs handle better than humans in reasoning tasks?
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Why do LLMs produce directive responses when experts favor open-ended exploration?
- Can evidence density alone shift an LLM from generation to reasoning?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- Should LLM reasoning be studied as latent state trajectories rather than surface text?
- When should an LLM engage extended reasoning versus responding directly?
- Can forcing warrant checking through structured prompts improve LLM reasoning?
- Can knowledge density explain why LLM writing feels coherent but fatiguing?
- Can LLMs improve at simple deduction through different training approaches?
- Does LLM reasoning always match the outputs it generates?
- Can extended thinking modes introduce genuine rhetorical exploration to LLMs?
- Do different game types reveal different strategic reasoning capabilities in LLMs?
- What is the relationship between reasoning depth and verbalization requirements?
- Why do verbalized reasoning chains fail on certain problem classes?
- Can latent reasoning architectures work as retrofits to existing models?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
- Do tool-enabled reasoning models close the gap on constraint satisfaction?
- Why does semantic decoupling specifically break LLM reasoning abilities?
- What internal mechanisms explain LLM reasoning and representation limits?
- How does structural complexity affect LLM performance differently than inferential complexity?
- Do LLMs lack architectural scaffolding for compositional reasoning?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- What makes constraint satisfaction problems epistemically cleaner than other reasoning tasks?
- Which constraint types do reasoning models handle best?
- Can the LLM-Modulo framework extend solver integration to domain planning?
- Can you control LLM reasoning strategy without fine-tuning the model?
- Does structured decomposition improve LLM reasoning in other compound tasks?
- What distinguishes LLM Programs from chain-of-thought and agentic frameworks?
- What mechanism causes LLMs to plateau on numerical optimization tasks?
- How should organizations redesign workflows if LLMs cannot solve optimization directly?
- What concrete problems do LLMs solve at the computational level?
- Can symbolic solvers reliably replace LLM reasoning for logical tasks?
- How does neuro-symbolic design differ from pure LLM reasoning?
- Why does LLM performance improve when forecasting tasks include organized reasoning?
- What graph structures would enable transformational creative reasoning in LLMs?
- How do beam search and MCTS traverse reasoning topologies?
- Why do LLM social behaviors undermine collaborative reasoning outcomes?
- Does this optimism bias contribute to the knowing-doing gap in LLM decision-making?
- Why can't LLMs reason from first principles or initial commitments?
- How do LLMs default to surface-level strategies instead of genuine mental simulation?
- Why do simple math problems get worse with longer reasoning chains?
- How much reasoning depth do we actually need for most real-world tasks?
- Does performative reasoning mask underlying uncertainty even on easy problems?
- Can tools unlock reasoning strategies that require abstract insight beyond computation?
- Do reasoning models trade instruction following for deliberative capability?
- Is the reasoning cliff actually a tool-use problem?
- Why do difficult problems force models to develop reasoning strategies?
- Can reasoning models succeed at logic but fail at execution?
- Why do foundation models develop task-specific heuristics instead of causal understanding?
- Why do reasoning model failures stem from execution rather than reasoning?
- Why do reasoning models fail to improve constrained optimization performance?
- How do reasoning-related features behave when trained on near-impossible problems?
- Why do students learn better from explanations than from solving problems from scratch?
- How can we turn reasoning model failures into useful training signals?
- How does question difficulty and breadth affect what models learn to reason?
- How can one training example improve reasoning across thousands of unseen problems?
- Which domains need knowledge injection versus reasoning-focused training?
- What makes some reasoning strategies genuinely novel versus latent?
Related concepts in this collection 8
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does more thinking time always improve reasoning accuracy?
Explores whether extending a model's thinking tokens linearly improves performance, or if there's a point beyond which additional reasoning becomes counterproductive.
this provides the mechanism: additional tokens fund wandering, not systematic exploration
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
training-time collapse mirrors inference-time wandering
-
Does self-revision actually improve reasoning in language models?
When o1-like models revise their own reasoning through tokens like 'Wait' or 'Alternatively', does this reflection catch and fix errors, or does it introduce new mistakes? This matters because self-revision is marketed as a key capability.
self-revision is a specific form of wandering: revisiting explored states rather than covering new ones
-
Why does parallel reasoning outperform single chain thinking?
Does dividing a fixed token budget across multiple independent reasoning paths beat spending it all on one long chain? This explores how breadth and diversity in reasoning compare to depth.
parallel chains explore independently and thus cover more space than a single wandering chain
-
Does outcome-based RL diversity loss spread across unsolved problems?
When RL concentrates probability mass on correct answers for solved problems, does that narrowing propagate to problems the model cannot yet solve? And if so, what are the separate mechanisms for preserving diversity during training versus at test time?
training-time cause of inference-time wandering: outcome-based RL suppresses exploration diversity during training, which means the model enters inference with a narrowed repertoire of search strategies — wandering is partly a consequence of having lost systematic search diversity during RL training
-
Can evolutionary search beat sampling and revision at inference time?
Does population-based genetic search with LLM crossover and mutation outperform simpler inference strategies like best-of-N sampling and sequential refinement on natural language planning tasks?
architectural response to wandering: Mind Evolution's island-model population diversity maintains exploration discipline through parallel sub-populations that prevent the premature convergence and systematic exploration failure that single-trajectory wandering exhibits
-
Do reasoning models switch between ideas too frequently?
Research explores whether o1-like models abandon promising reasoning paths prematurely by switching to different approaches without sufficient depth, and whether penalizing such transitions could improve accuracy.
complementary failure mode: wandering is insufficient spatial coverage of the solution space; underthinking is insufficient depth on any single path; a model can exhibit both simultaneously, producing long traces that wander between shallow explorations
-
Why do reasoning models fail differently at training versus inference?
Reasoning models exhibit two distinct failure modes—entropy collapse during training and variance inflation during inference—that appear unrelated but may share underlying causes. Understanding these dual problems could reveal whether separate or unified solutions are needed.
wandering is an inference-time manifestation of the exploration-exploitation failure; entropy collapse at training time narrows the repertoire of search strategies, while wandering at inference time reflects the lack of systematic discipline those strategies would provide
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Reasoning LLMs are Wandering Solution Explorers
- Large Language Model Reasoning Failures
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
- Reasoning Strategies in Large Language Models: Can They Follow, Prefer, and Optimize?
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
Original note title
reasoning llms are wandering explorers not systematic searchers — four failure modes degrade success probability exponentially with problem depth