SYNTHESIS NOTE
Topics›Reasoning Logic Internal Rules›this note

Do large language models reason symbolically or semantically?

Can LLMs follow explicit logical rules when those rules contradict their training knowledge? Testing whether reasoning operates independently of semantic associations reveals what computational mechanisms actually drive LLM multi-step inference.

Synthesis note · 2026-02-22 · sourced from Reasoning Logic Internal Rules

The "In-Context Semantic Reasoners" paper tests a fundamental question about what drives LLM reasoning by systematically decoupling semantics from the reasoning process across deduction, induction, and abduction tasks. The findings are clear: when semantics are consistent with commonsense, LLMs perform well; when semantics are removed or made counter-commonsense, performance collapses even when correct rules are provided in context.

The experimental design is precise. By replacing relation labels with shuffled alternatives ("motherOf" → "sisterOf", "female" → "male"), the researchers create tasks where the in-context rules are logically valid but semantically counter-intuitive. LLMs cannot follow these counter-commonsense rules despite having them explicitly in the prompt. The model's parametric knowledge — its compressed commonsense from training — overrides the in-context logical structure.

This reveals a specific computational mechanism: LLMs create "superficial logical chains" through semantic token associations, not through symbolic manipulation. The connections between tokens that enable multi-step reasoning are semantic connections, not logical ones. When those semantic connections support the correct answer, reasoning appears to work. When they conflict, reasoning fails regardless of what the prompt says.

The implication is that LLM reasoning is fundamentally bounded by training distribution semantics. Since Can large language models translate natural language to logic faithfully?, the failure is bidirectional: LLMs can neither translate TO formal logic faithfully nor reason FROM formal logic when it conflicts with semantic priors. Since Do foundation models learn world models or task-specific shortcuts?, the semantic dependency IS the heuristic — the model uses semantic similarity as a proxy for logical validity.

This connects to the Dual Process Theory framework: human System II symbolic reasoning operates independently of semantic content, but LLM "reasoning" remains entangled with System I semantic associations. The paper's suggestion — integrating LLMs with external non-parametric knowledge bases and improving in-context knowledge processing — implicitly acknowledges that the LLM alone cannot escape this limitation.

Retort implication — rules out a class of anthropomorphization: The finding constrains what we can say about LLM behavior in other domains. Any account that treats LLMs as agents who "reverse-engineer" justifications for conclusions they have committed to — the standard anthropomorphization of sycophancy, rationalization, or motivated reasoning — presupposes the semantic competence this note shows LLMs lack. If reasoning collapses when semantics are decoupled, there is no separable reasoning faculty available to perform a post-hoc rationalization. What looks like reverse-engineering is pattern-matching within semantic associations. This rules out a whole class of AI commentary that treats LLMs as dishonest agents who could have reasoned correctly but chose not to.

Metaphor as paradigmatic semantic decoupling: Metaphor is the literary instantiation of this finding. A metaphor works by using one domain's vocabulary to illuminate another — "time is money," "argument is war," "memory is a jar of flies." The decoupling between the source domain's semantics and the target domain's meaning is the defining feature of metaphorical language. Since LLM reasoning collapses when semantics are decoupled from their typical packaging, and metaphor is decoupled semantics, this predicts a specific failure mode: LLMs should handle conventional metaphors (lexicalized, semantically consistent with commonsense) better than novel literary metaphors (where the mapping between domains is unexpected and requires conceptual reasoning beyond semantic association). The Diplomat dataset (Diplomat: A Dialogue Dataset for Situated PragMATic Reasoning) suggests treating all figurative language as a unified pragmatic reasoning task — but the semantic-decoupling finding predicts that this unified approach will hit a wall at the novelty threshold where metaphors stop relying on conventional semantic associations.

Inquiring lines that read this note 264

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why don't LLMs reliably translate capability into accurate outputs? Can prompt-based context override biases that were embedded during pretraining? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? How should designers communicate what AI systems truly are and can do? Is language model reasoning authentic and what causes models to reason? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? What structural properties of attention create systematic model biases? Do reasoning traces faithfully reflect actual model reasoning? What causes reasoning models to fail or wander off track? How do neural networks achieve compositional generalization at scale? What compositional reasoning failures limit large language models despite scale? Do language models learn genuine understanding or just surface patterns? Can models improve accuracy without degrading reasoning quality? How effectively can language models perform reasoning, especially combined with symbolic methods? Do language models reason through causal mechanisms or semantic associations? Does encoded knowledge in language models actually influence their outputs? Can diffusion models match autoregressive performance on language generation tasks? How do false presuppositions and sycophancy drive persistent false beliefs in models? Can language models build genuine grounding through interaction? Is reasoning capability latent in base models or created by post-training? Why do embedding systems fail to capture task-relevant relationships? How do surface patterns enable correct outputs but reduce robustness? What enables genuine semantic understanding in language models? Why do stronger reasoning capabilities create tradeoffs with instruction following? How should inference compute be allocated based on problem difficulty? How should retrieval systems handle complex multi-step reasoning? Can reasoning scale in latent space without tokens? What reasoning architectures enable models to solve complex problems efficiently? How much do training data properties shape model reasoning? Do language models reason like humans or mimic surface patterns? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? What makes distillation transfer some model capabilities while suppressing others? How much does training format versus domain influence reasoning? What training dynamics and scale trigger emergence of reasoning capabilities? Why do token-level mechanisms matter for learning to reason? Does transformer attention architecture inherently drive sycophancy? Why is hallucination an inevitable limitation of current language models? Can inference-time compute effectively substitute for model scale? Can intelligent routing over smaller models outperform scaling a single large model? What is the relationship between thinking tokens and reasoning accuracy? How does persona conditioning amplify demographic stereotyping and bias in models? How do multi-agent LLM systems fail distinctly compared to single agents? Do reasoning benchmarks predict model performance in long-horizon workflows? Why do some clarifying approaches produce understanding while others just satisfy? Why does adding new knowledge through fine-tuning degrade existing capabilities? Does RL create genuinely new reasoning capabilities or refine existing ones? How do soft reasoning mechanisms explore multiple paths without explicit training? What types of diversity prevent reasoning systems from collapsing? Do language models possess genuine introspective self-awareness or only behavioral mimicry? What safeguards enable trustworthy AI-assisted scientific peer review at scale? How does reasoning length affect model performance across different tasks? How do prompt design choices influence model reasoning and performance? Can mechanistic interpretability reliably guide practical model design choices? How should systems decide whether to retrieve or reason alone? What capability trade-offs arise from domain specialization through fine-tuning? How do LLM judges' systematic biases affect alignment and evaluation outcomes?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 180 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms are in-context semantic reasoners not symbolic reasoners — when semantics are decoupled reasoning collapses