SYNTHESIS NOTE
Topics›MechInterp›this note

Do language models understand in fundamentally different ways?

Does mechanistic evidence reveal distinct tiers of understanding in LLMs—from concept recognition to factual knowledge to principled reasoning? And do these tiers coexist rather than replace each other?

Synthesis note · 2026-04-18 · sourced from MechInterp

This paper synthesizes mechanistic interpretability findings into a philosophical framework that moves beyond the binary "does AI understand?" debate. The framework proposes three hierarchical tiers:

Tier 1: Conceptual understanding — arises when a model forms "features" as directions in latent space that unify diverse manifestations of a single entity or property. This is the representational foundation: the model has learned that different surface forms connect to the same underlying concept. MI evidence: SAE features, linear probing, representation geometry studies all demonstrate this.

Tier 2: State-of-the-world understanding — arises when the model learns contingent factual connections between features and dynamically tracks changes. "Michael Jordan is a basketball player" is not just a high-probability string but a reflection of an internal model linking the Michael Jordan concept to the basketball player concept. This goes beyond association to structured knowledge representation.

Tier 3: Principled understanding — arises when the model discovers compact "circuits" that connect facts via general rules rather than memorizing each fact individually. This is the shift from knowing that to knowing why. The grokking literature provides the clearest evidence: models that transition from memorization to generalization develop circuits implementing actual algorithmic rules (e.g., modular addition via Fourier transforms).

The critical insight is that higher-tier mechanisms coexist with lower-tier heuristics rather than replacing them. A model can have principled understanding of arithmetic in one circuit while relying on pattern-matching heuristics in another. This heterogeneity means understanding is not a single binary property but a patchwork: principled in some domains, merely conceptual in others, and purely heuristic in yet others.

This has direct implications for trust and deployment. The fact that a model demonstrates principled understanding in one domain gives no guarantee that it operates at the same tier in adjacent domains. The coexistence of understanding tiers also explains why models can be simultaneously impressive and brittle: the principled circuits work reliably, but the heuristic patches fail unpredictably.

Inquiring lines that read this note 79

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can mechanistic interpretability reliably guide practical model design choices? Why don't LLMs reliably translate capability into accurate outputs? Does encoded knowledge in language models actually influence their outputs? How effectively can language models perform reasoning, especially combined with symbolic methods? Can language models build genuine grounding through interaction? What causes reasoning models to fail or wander off track? Is language model reasoning authentic and what causes models to reason? Do language models possess genuine introspective self-awareness or only behavioral mimicry? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What enables genuine semantic understanding in language models? Do language models reason like humans or mimic surface patterns? Do language models learn genuine understanding or just surface patterns? What design and behavioral factors drive false consciousness attribution to AI? Do language models reason through causal mechanisms or semantic associations? How can we distinguish genuine model deception from honest errors? Why do stronger reasoning capabilities create tradeoffs with instruction following? What training dynamics and scale trigger emergence of reasoning capabilities? How should systems decide whether to retrieve or reason alone? How do neural networks achieve compositional generalization at scale? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Is reasoning capability latent in base models or created by post-training? Why doesn't reasoning volume improve theory of mind performance? What compositional reasoning failures limit large language models despite scale?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 150 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

mechanistic interpretability evidence supports three hierarchical varieties of LLM understanding — conceptual then state-of-world then principled — each tied to a distinct computational organization