Do large language models reason symbolically or semantically?
Can LLMs follow explicit logical rules when those rules contradict their training knowledge? Testing whether reasoning operates independently of semantic associations reveals what computational mechanisms actually drive LLM multi-step inference.
The "In-Context Semantic Reasoners" paper tests a fundamental question about what drives LLM reasoning by systematically decoupling semantics from the reasoning process across deduction, induction, and abduction tasks. The findings are clear: when semantics are consistent with commonsense, LLMs perform well; when semantics are removed or made counter-commonsense, performance collapses even when correct rules are provided in context.
The experimental design is precise. By replacing relation labels with shuffled alternatives ("motherOf" → "sisterOf", "female" → "male"), the researchers create tasks where the in-context rules are logically valid but semantically counter-intuitive. LLMs cannot follow these counter-commonsense rules despite having them explicitly in the prompt. The model's parametric knowledge — its compressed commonsense from training — overrides the in-context logical structure.
This reveals a specific computational mechanism: LLMs create "superficial logical chains" through semantic token associations, not through symbolic manipulation. The connections between tokens that enable multi-step reasoning are semantic connections, not logical ones. When those semantic connections support the correct answer, reasoning appears to work. When they conflict, reasoning fails regardless of what the prompt says.
The implication is that LLM reasoning is fundamentally bounded by training distribution semantics. Since Can large language models translate natural language to logic faithfully?, the failure is bidirectional: LLMs can neither translate TO formal logic faithfully nor reason FROM formal logic when it conflicts with semantic priors. Since Do foundation models learn world models or task-specific shortcuts?, the semantic dependency IS the heuristic — the model uses semantic similarity as a proxy for logical validity.
This connects to the Dual Process Theory framework: human System II symbolic reasoning operates independently of semantic content, but LLM "reasoning" remains entangled with System I semantic associations. The paper's suggestion — integrating LLMs with external non-parametric knowledge bases and improving in-context knowledge processing — implicitly acknowledges that the LLM alone cannot escape this limitation.
Retort implication — rules out a class of anthropomorphization: The finding constrains what we can say about LLM behavior in other domains. Any account that treats LLMs as agents who "reverse-engineer" justifications for conclusions they have committed to — the standard anthropomorphization of sycophancy, rationalization, or motivated reasoning — presupposes the semantic competence this note shows LLMs lack. If reasoning collapses when semantics are decoupled, there is no separable reasoning faculty available to perform a post-hoc rationalization. What looks like reverse-engineering is pattern-matching within semantic associations. This rules out a whole class of AI commentary that treats LLMs as dishonest agents who could have reasoned correctly but chose not to.
Metaphor as paradigmatic semantic decoupling: Metaphor is the literary instantiation of this finding. A metaphor works by using one domain's vocabulary to illuminate another — "time is money," "argument is war," "memory is a jar of flies." The decoupling between the source domain's semantics and the target domain's meaning is the defining feature of metaphorical language. Since LLM reasoning collapses when semantics are decoupled from their typical packaging, and metaphor is decoupled semantics, this predicts a specific failure mode: LLMs should handle conventional metaphors (lexicalized, semantically consistent with commonsense) better than novel literary metaphors (where the mapping between domains is unexpected and requires conceptual reasoning beyond semantic association). The Diplomat dataset (Diplomat: A Dialogue Dataset for Situated PragMATic Reasoning) suggests treating all figurative language as a unified pragmatic reasoning task — but the semantic-decoupling finding predicts that this unified approach will hit a wall at the novelty threshold where metaphors stop relying on conventional semantic associations.
Inquiring lines that read this note 264
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations?- Can LLMs infer situational context the way humans do pragmatically?
- Can explicit connectives compensate for missing intentional tracking in LLMs?
- Can LLMs improve at metaphor if they handle decoupled semantics better?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Why do explicit discourse connectives help LLMs but implicit relations cause failures?
- How does the symbol grounding problem apply to artificial language systems?
- Can LLMs infer implicit meaning without surface linguistic markers?
- How do embedding contexts like presupposition triggers affect LLM entailment reasoning?
- Can LLMs identify implicit metaphoric mappings that require pragmatic inference?
- Why do LLMs choose surface-order quantifier scope over contextually correct readings?
- Can LLM semantic representations exist without causally influencing their generation output?
- Why do LLMs perform better on explicit discourse connectives than implicit relations?
- Can LLMs compute how presuppositions project through embedded clauses?
- Why do LLMs struggle to translate natural language into logical formalizations?
- Can we use LLM language without adopting LLM assumptions?
- What semantic information is necessary to preserve for sound LLM reasoning?
- Do language models need words to think or just latent structure?
- Why do LLMs fall for and deploy logical fallacies with equal confidence?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- How do LLMs compress specific expert knowledge into median abstraction?
- Which knowledge types do LLMs handle better than humans in reasoning tasks?
- Why do LLMs fail at iterative numerical computation in latent space?
- How faithful are natural language explanations from LLMs really?
- What levels of understanding about LLM knowledge representation can automated systems reliably extract?
- How does surface salience compete with background knowledge in model inference?
- Can prompt-based debiasing overcome entrenched LLM model priors?
- Can explicit numerical signals override learned linguistic defaults in fine-tuned models?
- How does inductive reasoning from partial evidence enable hypothesis formation?
- Why does monological training prevent models from overriding statistical priors?
- How do training associations override context information in language models?
- How does LLM-PKG compare to mining product relations directly from interaction data?
- How do LLMs and knowledge graphs work together in different integration patterns?
- Can explicit linkers replace vector similarity for multi-step question answering?
- Can knowledge graph structure alone generate sufficient training signals for domain reasoning?
- Why do LLMs recognize graph entities without modeling their relationships?
- Can evidence density alone shift an LLM from generation to reasoning?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- Should LLM reasoning be studied as latent state trajectories rather than surface text?
- Why does hypothesis attestation bias exist separately from frequency bias in NLI?
- Do LLMs understand implicit warrants in reasoning chains?
- Why can LLMs identify argument structure but not check warrants?
- Why do LLMs fail when asked to use counter-commonsense rules explicitly?
- What specific linguistic features cause LLMs to fail at trivial entailment?
- Can LLMs improve at simple deduction through different training approaches?
- Does LLM reasoning always match the outputs it generates?
- Why do LLMs fail at counterfactual reasoning despite factual knowledge?
- Why do smaller LLMs fail at zero-shot argument scheme classification?
- Does compressing Walton's schemes into nine categories make LLM classification easier?
- Can LLM-generated descriptions of schemes outperform formal dictionary definitions for prompting?
- Why do LLM descriptions of argument schemes work better than formal definitions for classification?
- Can irrelevant information reliably expose the limits of LLM reasoning?
- How do different LLMs converge on similar argumentative structures independently?
- Can LLM reasoning traces be validated against actual population reasoning?
- How do transformers perform multi-hop reasoning across distant training documents?
- Can symbolic mechanisms improve transformer compositional abilities?
- Can explicit stack mechanisms extend what formal languages transformers can learn?
- Can transformers abstract relational structure without explicit symbolic machinery?
- Why do simple length heuristics outperform sophisticated semantic methods?
- Why do language model reasoning chains look fluent when they deviate from the task?
- How does the frame problem differ between symbolic and statistical reasoning systems?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- How do humans and LMs differ on multi-hop reasoning?
- Where do humans and language models actually diverge in reasoning ability?
- Why does explicit reasoning degrade passage reranking performance?
- What causes snowball errors to accumulate across reasoning steps in language models?
- Can long-context models handle compositional reasoning requiring structured logic?
- Why do reasoning models fail when input length increases even below context limits?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- Why does cross-text analogical reasoning fail when semantics decouple from symbols?
- Why does removing semantic content collapse reasoning in language models?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- Why do smaller models lose reasoning faithfulness more than larger models?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- What makes hierarchical reasoning effective for taxonomy induction?
- Can neural networks represent symbolic structures without explicit mechanisms?
- Can neural networks learn that A implies B in reverse?
- What non-parametric methods could replace latent factors for inductive learning?
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- How should we rethink the symbolism versus connectionism debate in light of LLMs?
- What prevents LLM representations from causally influencing generation outputs?
- How does syntactic encoding relate to semantic feature representation?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Why do true and false LLM outputs use the same mechanism?
- Why do language models struggle with formal logical reasoning and joins?
- Why does augmenting natural language with formal representations outperform full formalization?
- What other structural limits exist at the language-formal boundary?
- Why do long-context language models struggle with compositional reasoning tasks?
- Can language models translate theorems faithfully without semantic loss?
- Do language models learn surface patterns instead of underlying linguistic principles?
- Can language models reason without relying on learned semantic patterns?
- Why do language models imitate reasoning form without abstract inference capability?
- Do LLMs compute scalar implicature differently across conversational contexts?
- Does generalization frequency explain why models favor upward semantic movement?
- Do LLMs rely on surface heuristics instead of learning recursive grammar rules?
- Can complexity-stratified testing reveal whether LLMs understand grammatical structure?
- How does structural depth in sentences predict LLM annotation accuracy?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Can language models reason without relying on surface level pattern matching?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- What empirical evidence supports the Learning Law on real language models?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Why does compositional reasoning fail to explain cross-domain transfer?
- How does the outer loop escape its own LLM's knowledge boundaries when discovering mechanisms?
- Can symbolic solvers rescue language models from logical reasoning failures?
- Can reasoning chains work without logical validity?
- Do reasoning languages like Prolog follow the same two-constraint transfer pattern?
- What makes symbolic operations different from general knowledge questions?
- Why does semantic decoupling specifically break LLM reasoning abilities?
- Why do LLMs generate logical forms without preserving semantic content?
- Which game type reveals minimax reasoning in language models?
- What internal mechanisms explain LLM reasoning and representation limits?
- Can LLMs translate between natural language and formal logic faithfully?
- How does context complexity affect LLM performance on temporal reasoning tasks?
- Why can LLMs interpret formal logic better than they generate it?
- How does structural complexity affect LLM performance differently than inferential complexity?
- How does semantic reasoning differ from symbolic reasoning in language models?
- Can LLMs reliably generate novel working architectures without structured representations?
- Do LLMs lack architectural scaffolding for compositional reasoning?
- What makes deductive reasoning so brittle in language models overall?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- What makes structural logic correlate so strongly with contextual consistency?
- How does in-context semantic reasoning differ from symbolic reasoning in concept fusion?
- Can language models perform purely symbolic reasoning when semantics are removed?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Can language models perform genuine symbolic reasoning without semantic grounding?
- Can you control LLM reasoning strategy without fine-tuning the model?
- Does structured decomposition improve LLM reasoning in other compound tasks?
- What latent mechanisms do LLMs use when they cannot execute iterative methods?
- How do deterministic symbolic solvers improve the reliability of language model reasoning?
- How do LLMs translate informal prose into logically correct formal specifications?
- How do LLMs lose information when translating natural language to formal logic?
- Why do LLMs fail at faithful autoformalisation of reasoning problems?
- Can symbolic solvers reliably replace LLM reasoning for logical tasks?
- What makes natural language reasoning more practical than formal languages for multi-framework codebases?
- How does neuro-symbolic design differ from pure LLM reasoning?
- Can LLMs simultaneously reason and optimize their own modules?
- How should LLM abstraction tools be evaluated without manual labeling?
- How do humans use associative reasoning without causal connections?
- Why do LLMs inherit causal biases from their training data?
- Do LLMs rely on surface statistical patterns instead of causal structure?
- Why does LLM compression eliminate causal grounding in conceptual representations?
- Can LLMs reason through semantics without understanding causal mechanisms?
- How does semantic association differ from mechanistic causal reasoning?
- Why do LLMs reason fluently about causality but lack causal rigor?
- Do LLMs show stronger reasoning about causality than about temporal ordering?
- How deeply are ideological structures represented in large language models?
- Can we distinguish between semantic and symbolic reasoning in language models?
- Why do language models substitute parametric knowledge over retrieved context mid-reasoning?
- Is relevant knowledge encoded in LMs but not causally active in generation?
- What reveals the epistemic limits of language models?
- Why do explicit linguistic markers override semantic computation in models?
- What sparse mechanistic structures drive reasoning traces in language models?
- Why do language models produce unfaithful chain of thought explanations?
- What implicit premises do language models skip even with correct surface reasoning?
- How do pretrained language models represent inferential patterns versus lexical and positional cues?
- What geometric structure do language models actually use during inference?
- How do semantic and symbolic reasoning capabilities differ in language models?
- How does tool-based reasoning expand what language models can do?
- What are the stages of inference inside language models?
- Can autoregressive models learn faithful translation to logical representations without semantic loss?
- Why do diffusion LLM answer tokens converge in confidence long before reasoning stabilizes?
- What circuit mechanisms produce belief bias in syllogistic reasoning?
- Why do non-factive verbs and triggers both fool language models?
- How does semantic grounding differ between human minds and language models?
- Can training LLMs to form ad-hoc conventions improve their pragmatic reasoning?
- How does business logic specification replace annotated training datasets?
- Can targeted activation steering surface latent reasoning in base models?
- What makes reasoning-specific post-training different from standard parameter scaling?
- Why do recursive belief models require different training than logical derivation?
- How do single training examples activate reasoning capabilities in language models?
- Do base models contain latent reasoning that minimal training can unlock?
- Do base models truly possess latent reasoning capability?
- Does latent reasoning capability exist in base models before any training?
- Can models reason at inference without specialized internal training?
- How much training data is truly necessary to unlock latent model reasoning?
- Does the base model already contain latent reasoning capability?
- What does pass@k reveal about base model reasoning capacity?
- Can models possess latent reasoning capability that training signals fail to unlock?
- What mechanisms activate latent reasoning capabilities already present in base models?
- Can minimal training signals unlock latent reasoning capability in base models?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
- What latent reasoning capability do base models already possess before training?
- Why do embeddings measure semantic association instead of task relevance?
- Why do unit-sphere spaces fail at distinguishing word order and negation?
- Can language models acquire meaning from distributional patterns alone without joint attention?
- Do metaphors work by decoupling meaning from linguistic associations?
- What separates pattern matching from genuine language understanding?
- How does bidirectional entailment distinguish semantic equivalence from token similarity?
- How do corpus statistics shape the abstraction hierarchy in language model representations?
- Why do models fail on logically equivalent tasks with different data distributions?
- Why do models learn reasoning form instead of actual abstract inference?
- Why do reasoning-optimized models show no sycophancy resistance advantage?
- Can machine learning encode pragmatic reasoning about when rules should bend?
- Why do reasoning-optimized models show no resistance advantage on agreement tasks?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
- Do latent sequence vectors outperform per-token latent iterative computation for reasoning?
- How do LLMs infer information that was explicitly censored?
- Can latent reasoning mechanisms and recursive tracking mechanisms be combined effectively?
- How do recursive language models rethink where to store reasoning?
- Can continuous latent reasoning match discrete chain-of-thought without training modifications?
- Can latent reasoning achieve the same substitution without tokens?
- Can structured workflows unlock latent reasoning abilities that raw models don't show?
- Do models leak their true associations through reasoning traces and behavior?
- Why do rare complex structures in training data harm LLM generalization?
- How much does training composition affect syntactic versus reasoning performance?
- How does training data distribution constrain LLM moral reasoning patterns?
- Why can't LLMs reason from first principles or initial commitments?
- Can LLMs simulate belief revision in social systems without modeling thought?
- Why don't LLM explanations predict what models would actually do?
- How do explicit reasoning traces help models construct valid syntactic trees?
- How do we verify that stated beliefs actually follow from underlying motifs?
- Can instance-adaptive reasoning happen without sequential token dependencies?
- Why does unstructured chain-of-thought permit assumption-based errors that templates prevent?
- How does latent reasoning recursion compare to chain-of-thought reasoning?
- Is verbalized chain-of-thought necessary for language model reasoning?
- How does training data format shape whether models reason in parallel or sequentially?
- Does training data format shape which reasoning strategies LLMs develop?
- How does an instruction-following LLM activate latent retrieval knowledge?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- Do reflection tokens and symbolic tokens serve different roles in reasoning?
- Why does hierarchical formal language training improve token efficiency more than natural language?
- What evidence shows that reasoning chains encode token-level functional structure?
- Can standard next-token prediction capture complex multi-step human reasoning directly?
- Where does inference compute stop substituting for model capacity?
- Can non-variational posterior approximation schemes deliver comparable reasoning improvements?
- Can reinforcement learning close the gap between LLM reasoning and action?
- Can base models spontaneously produce reasoning traces without any RL training?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can large language models translate natural language to logic faithfully?
This explores whether LLMs can convert natural language statements into formal logical representations without losing meaning. It matters because faithful translation is essential for any AI system that reasons formally or verifies specifications.
bidirectional semantic dependency: fails translating TO logic and reasoning FROM logic
-
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
semantic associations are the heuristic mechanism
-
Why do language models ignore information in their context?
Explores why language models sometimes override contextual information with prior training associations, and whether providing more context can solve this problem.
same mechanism: parametric knowledge overrides in-context information
-
Does semantic grounding in language models come in degrees?
Rather than asking whether LLMs truly understand meaning, this explores whether grounding is actually a multi-dimensional spectrum. The question matters because it reframes the sterile understand/don't-understand debate into measurable, distinct capacities.
functional grounding through semantic associations explains why reasoning works within commonsense boundaries
-
Why do neural networks fail at compositional generalization?
Exploring whether the binding problem from neuroscience explains neural networks' inability to systematically generalize. The binding problem has three aspects—segregation, representation, and composition—each creating distinct failure modes in how networks handle structured information.
the binding problem may explain WHY semantic decoupling collapses reasoning: without compositional binding mechanisms, removing semantic content removes the only glue holding multi-step inference together; semantic associations serve as a substitute for genuine compositional binding
-
Do LLMs actually have world models or just facts?
The term 'world model' conflates two different capabilities: factual representation versus mechanistic understanding. Understanding which one LLMs actually possess matters for assessing their reasoning reliability.
semantic reasoning operates on factual world representation (Sense 1) but cannot perform mechanistic reasoning (Sense 2) when logic must override semantic priors
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
- Language models show human-like content effects on reasoning tasks
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
- Can Large Language Models Reason and Optimize Under Constraints?
Original note title
llms are in-context semantic reasoners not symbolic reasoners — when semantics are decoupled reasoning collapses