Can language models understand without actually executing correctly?
Do LLMs truly comprehend problem-solving principles if they consistently fail to apply them? This explores whether the gap between articulate explanations and failed actions points to a fundamental architectural limitation.
LLMs display surface fluency yet systematically fail at tasks requiring symbolic reasoning, arithmetic accuracy, and logical consistency. The diagnosis: a persistent gap between comprehension and competence, rooted not in knowledge access but in computational execution.
The paper names this "computational split-brain syndrome" — instruction and action pathways are geometrically and functionally dissociated within the model. The model can articulate the correct principle for how to solve a problem, then fail to apply that principle in the next step. This is not forgetting, not hallucination, not knowledge deficit — it is a structural disconnect between knowing-how-to-describe and knowing-how-to-do.
The failure recurs across domains: mathematical operations, relational inferences, logical deductions. The consistency across domains suggests an architectural rather than domain-specific cause. LLMs function as powerful pattern completion engines but lack the scaffolding for principled, compositional reasoning — structure for executing what they can describe.
This provides a mechanistic name for Can LLMs understand concepts they cannot apply?. Potemkin understanding names the phenomenon; computational split-brain names the mechanism. The geometric separation between instruction representations and execution pathways explains why the model can generate correct explanations and incorrect applications simultaneously without detecting the inconsistency.
It also concretizes Why do language models fail to act on their own reasoning?. The 87% vs 64% gap is the quantitative signature of the split-brain: the instruction pathway (rationale generation) and the execution pathway (action selection) draw on overlapping but dissociated representations.
The paper further argues that mechanistic interpretability findings may reflect training-specific pattern coordination rather than universal computational principles — the internal structures we discover may be execution artifacts, not reasoning architecture.
Planning as the paradigmatic test case. The 8-puzzle study (On the Limits of Innate Planning in Large Language Models) isolates two specific deficits: (1) brittle internal state representations leading to frequent invalid moves, and (2) weak heuristic planning with models entering loops or selecting actions that don't reduce distance to the goal. Even with an external move validator providing only valid moves, none of the models solve any puzzles. The comprehension-competence split is stark: models can articulate puzzle-solving strategies but cannot maintain accurate state representations across sequential moves. Since Can large language models actually create executable plans?, the gap widens with task complexity: 87% correct rationales → 64% correct actions → 12% executable plans → 0% puzzle solutions with validator assistance.
Inquiring lines that read this note 95
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why does adding new knowledge through fine-tuning degrade existing capabilities? Why don't LLMs reliably translate capability into accurate outputs?- Where do LLMs succeed at generation but struggle with evaluation?
- Why do LLMs generate ideas that sound novel but fail during execution?
- What specific execution barriers do LLM ideas encounter most frequently?
- What distinguishes planning knowledge from an executable plan that works?
- Why do LLM explanations feel authoritative even when alignment with the model fails?
- What explains the 87 percent to 12 percent cliff in plan executability?
- Why does LLM knowledge fail to influence their actual outputs?
- Can LLMs explain concepts correctly while failing to use them?
- What causes LLMs to ignore unstated constraints they know about?
- Why do LLMs excel at generation but struggle with evaluation?
- Where do LLMs fail as knowledge systems compared to humans?
- Why do LLMs explain evidence accurately while missing its implications?
- Do LLMs fail exploration because of context integration or computational limitations?
- How can a model explain something correctly yet fail to apply it?
- Why do benchmark tests fail to detect LLM comprehension gaps?
- Why do LLMs choose incorrect edits despite understanding the task?
- Can surface-level correctness hide failures in structural learning by LLMs?
- Can we systematically enumerate LLM failure modes from first principles?
- How faithful are natural language explanations from LLMs really?
- Can LLMs reliably audit other language models for errors?
- What barriers prevent experts from specifying concepts for LLM extraction?
- Why do people misinterpret or misuse LLM outputs in practice?
- What other latent LLM capabilities remain inactive without explicit activation cuing?
- What cognitive capacities do LLMs actually lack that commentary assumes they have?
- Can LLMs participate meaningfully in discourse without consciousness or understanding?
- What structural limits prevent LLMs from abstracting moral principles?
- Why do LLMs struggle to translate natural language into logical formalizations?
- Can we use LLM language without adopting LLM assumptions?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- Why do LLMs fail when asked to use counter-commonsense rules explicitly?
- Why do LLMs struggle with negation and exception handling?
- What makes a problem instance unfamiliar to a language model?
- Why do language models fail at planning despite understanding strategies?
- What architectural changes would let language models develop genuine functional competence?
- Why do LLMs understand efficient language but fail to produce it?
- Why do language models fail at understanding ambiguous or complex requirements?
- How do dependency errors propagate through incorrectly formalized definitions?
- Can models identify what information they are missing in underspecified problems?
- Can LLMs learn to ask clarifying questions instead of guessing?
- Why do LLM agents make promises without executing them?
- Why do LLMs explain correct reasoning but then choose greedy actions?
- How does the outer loop escape its own LLM's knowledge boundaries when discovering mechanisms?
- What internal mechanisms explain LLM reasoning and representation limits?
- Why can LLMs interpret formal logic better than they generate it?
- How does structural complexity affect LLM performance differently than inferential complexity?
- Do LLMs lack architectural scaffolding for compositional reasoning?
- What mechanism causes LLMs to plateau on numerical optimization tasks?
- How should organizations redesign workflows if LLMs cannot solve optimization directly?
- What concrete problems do LLMs solve at the computational level?
- What latent mechanisms do LLMs use when they cannot execute iterative methods?
- What prevents monolithic LLMs from coordinating decomposition with execution?
- Can LLMs simultaneously reason and optimize their own modules?
- How should LLM abstraction tools be evaluated without manual labeling?
- Why do monological explanations fail to transfer understanding compared to dialogical ones?
- How do dialogue acts and explanation moves interact to predict understanding success?
- At what complexity does LLM discourse failure become practically harmful?
- What distinguishes genuine understanding from correct output without coherent principles?
- What makes some interpretive postures stick while others fail to form?
- Why can't LLMs reason from first principles or initial commitments?
- How do knowing and doing diverge in LLM decision-making?
- What structural framework prevents LLM explanations from becoming just plausible fiction?
- How do LLM explanations diverge from actual internal reasoning?
- Why don't LLM explanations predict what models would actually do?
- How might human-LLM teams reinforce each other's causal reasoning mistakes?
- Can LLMs reason through semantics without understanding causal mechanisms?
- Why do LLMs reason fluently about causality but lack causal rigor?
- Can mechanistic interpretability explain explanation-execution disconnection?
- How does the knowing-doing gap relate to Potemkin understanding?
- Can interventions on model components prove mechanism without explaining encoding?
- What structural features enable agents to detect when understanding has broken down?
- How can correct explanations coexist with failed applications in AI?
- Can reasoning models succeed at logic but fail at execution?
- Why do reasoning model failures stem from execution rather than reasoning?
- Can models distinguish between logical impossibility and their own execution limits?
- Why do strong models struggle more with instruction following than mid-tier ones?
- How does belief-behavior inconsistency relate to instruction execution splits?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLMs understand concepts they cannot apply?
Explores whether large language models can correctly explain ideas while simultaneously failing to use them—and whether that combination reveals something fundamentally different from ordinary mistakes.
Potemkin understanding is the phenomenon; split-brain is the mechanism
-
Why do language models fail to act on their own reasoning?
LLMs produce correct explanations far more often than they produce correct actions. What causes this knowing-doing gap, and can training methods close it?
the quantitative signature of the comprehension-competence dissociation
-
Do language models actually use their encoded knowledge?
Probes can detect that LMs encode facts internally, but do those encoded facts causally influence what the model generates? This explores the gap between knowing and doing.
the encoding≠generation gap is the representational version of the same split
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Large Language Model Reasoning Failures
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- LLMs can implicitly learn from mistakes in-context
Original note title
comprehension without competence is a distinct LLM failure mode — instruction and execution pathways are dissociated