Can LLM understanding rely on just representation or causation alone?
Explores whether mechanistic interpretability of language models requires both mapping what is encoded (representational analysis) and testing if that encoding drives behavior (causal analysis), or whether either method suffices alone.
The implementation-level argument in Levels of Analysis for LLMs is that representational analysis and causal analysis are partners, not alternatives. Representational analysis maps what information a model encodes — which features, circuits, attention heads carry which signals. Causal analysis tests whether the information that is encoded actually drives behavior — through interventions, ablations, activation patches. Either method alone produces an incomplete account: a representation that is encoded but causally inert is a curiosity, and a causal effect with no representational characterization is unexplained.
The synergy matters because both methods can fool you alone. Representational analysis can identify features that correlate with behavior without showing they cause it — a classic confound. Causal analysis can demonstrate that intervening on some component changes behavior without telling you what that component encodes — the lesion shows damage but not function. The combination — representational analysis locates candidates, causal analysis tests their functional role — is what produces mechanistic claims rather than descriptive ones.
This has methodological consequences for interpretability research. Studies that report only feature visualizations or only activation patches contribute, but they do not close the loop. The convergent evidence comes from pairs: locate a candidate feature representationally, then verify it causally; identify a causal component, then map its representation. The literature on attention circuits, induction heads, and feature dictionaries has been moving toward this pairing.
For LLM understanding specifically, this template explains why some claimed "mechanisms" have not held up. They were representational without causal verification (a feature that looked like task encoding but did not drive task behavior) or causal without representational characterization (an intervention that mattered but described nothing). The discipline imported from cognitive neuroscience is to demand both.
Inquiring lines that read this note 81
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Is language model reasoning authentic and what causes models to reason? What causes reasoning models to fail or wander off track? How should designers communicate what AI systems truly are and can do? Can mechanistic interpretability reliably guide practical model design choices?- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- How much do mechanistic interpretability findings reflect true reasoning architecture?
- Are detection and identification of injections truly separable in neural circuits?
- Can mechanistic interpretability explain explanation-execution disconnection?
- Does causal intervention alone explain how neural mechanisms implement representations?
- Can attractor dynamics compete with input-based probing for characterizing model knowledge?
- How does mechanistic interpretability complement learning mechanics in explaining deep learning?
- Why do attention circuits need causal verification beyond feature visualization?
- What distinguishes a representational feature from a causally inert correlation?
- How do ablation studies reveal function without representational characterization?
- Can interventions on model components prove mechanism without explaining encoding?
- Can mechanistic interpretability tools decode the biases alignment training conceals?
- How do mechanistic features compare to natural language for interpretability?
- How do mechanistic interpretability tools help distinguish truthfulness from honesty?
- What makes representation engineering better than mechanistic interpretability for detecting hidden objectives?
- Why do feature visualizations alone fail to establish mechanistic claims?
- How do probe-based interventions in activation space compare to mechanistic interpretability approaches?
- Can mechanistic interpretability findings guide practical interventions in model design?
- What computational methods most reliably establish causal evidence in AI mechanism discovery?
- How does representation engineering compare to mechanistic interpretability for auditing?
- Do causal rules enforce robustness that statistical patterns alone cannot maintain?
- What makes causal belief networks more auditable than prompted personas?
- Can causal models be extended to include non-causal cognition?
- How do world models create indirect causal grounding without physical environment contact?
- Do LLMs rely on surface statistical patterns instead of causal structure?
- Are traditional cognitive theories missing interaction effects between mechanisms?
- Can LLMs reason through semantics without understanding causal mechanisms?
- How does semantic association differ from mechanistic causal reasoning?
- How does vehicle causality differ from content causality in physical systems?
- Can models be trained to hide causal influences in their explanations?
- Why do LLMs reason fluently about causality but lack causal rigor?
- Can a Reflect mechanism detect and revise failed causal predictions?
- How does causal structure avoid behaviorist limitations in LLM social simulation?
- What architectural changes would help LLMs distinguish causal relationships from temporal sequences?
- How does the outer loop escape its own LLM's knowledge boundaries when discovering mechanisms?
- What internal mechanisms explain LLM reasoning and representation limits?
- How should LLM abstraction tools be evaluated without manual labeling?
- What inductive bias would force models to learn Newtonian mechanics instead of shortcuts?
- Can steering vectors prove that representations are genuinely organized?
- Can geometric structure in representations exist without supporting functional mechanisms?
- How do classical mechanics and statistical mechanics provide methodological templates for learning theory?
- Can spectral eigenvector ordering serve as a model-agnostic interpretability probe?
- Can representation analysis methods detect complex features models compute with?
- How should we rethink the symbolism versus connectionism debate in light of LLMs?
- What prevents LLM representations from causally influencing generation outputs?
- How do we distinguish knowledge encoding from knowledge usage in models?
- How do mechanistic interpretability methods surface what models represent internally?
- How much introspective capability do safety mechanisms actively suppress in models?
- Can LLMs have minimal introspection through causal linkage to internal states?
- Do causal histories determine what mental states a system can instantiate?
- Can functional behavior alone capture what makes something a genuine belief?
- Can a perfect behavioral simulation constitute genuine understanding or experience?
- What consumption data would validate the limited-consumption model in production systems?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- What's the difference between representing world facts and generating world mechanisms?
- Does next-state prediction alone build mechanistic world models or just sophisticated interpolation?
- How do world models decompose between representation of facts versus generative mechanisms?
- How do delayed effects complicate causal attribution in agent systems?
- What distinguishes mechanical generation failures from deliberate behavioral withholding?
- What structural framework prevents LLM explanations from becoming just plausible fiction?
- What makes LLM behavior socially interpretable to human observers?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can cognitive science methods unlock how LLMs actually work?
Does Marr's three-level framework—developed to understand biological minds—offer interpretability researchers the structured methodology they need to decode opaque language models?
same paper, the framework
-
Can we predict where language models will fail?
Does characterizing the abstract computational problem an LLM solves—as a probability machine over sequences—let us predict which tasks it will struggle with systematically, before running experiments?
same paper, computational level companion
-
Can psychology methods reveal what alignment training conceals?
Do indirect cognitive psychology techniques like the IAT expose LLM associations that direct questioning misses because alignment training teaches models to filter verbal responses? This matters for evaluating whether models truly lack biases or simply hide them.
same paper, algorithmic level companion
-
Do language model reasoning drafts faithfully represent their actual computation?
If models externalize reasoning in thinking drafts before answering, does the draft accurately reflect their internal process? This matters for AI safety monitoring and error detection.
adjacent: dual-dimension methodology in CoT
-
Does sandbagging use a single residual stream axis?
Whether language models hide capability by writing deceptive intent to one axis in the residual stream that a later layer reads and acts on. Understanding the mechanism matters for designing targeted interventions.
an instance of the pairing: an axis for a behavior and a single-layer graft that tests it; the excerpt does not say how the axis was found, and it is shown only on installed locks
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Mechanistic Indicators of Understanding in Large Language Models
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Levels of Analysis for Large Language Models
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- Rethinking Large Language Models in Mental Health Applications
- Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis
- Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
Original note title
mechanistic understanding of LLMs requires both representational analysis and causal analysis — either alone is insufficient