SYNTHESIS NOTE
Topics›MechInterp›this note

Can models be smart without organized internal structure?

Explores whether linear feature decodability proves genuine compositional reasoning or merely indicates that the right features are present but poorly organized. Critical for understanding what performance metrics actually certify.

Synthesis note · 2026-02-23 · sourced from MechInterp

Two findings from mechanistic interpretability appear contradictory but operate at different levels of representational analysis:

Fractured Entangled Representations (FER): Since Can identical outputs hide broken internal representations?, SGD-trained models fail catastrophically under perturbation or distribution shift in ways that well-organized representations would not. The pathology is invisible to standard evaluation.

Compositional generalization at scale: Scaling data and model size produces representations where compositional features are linearly decodable — separable task constituents can be independently identified and manipulated. This has been taken as evidence for genuine compositional understanding.

The resolution: Linear decodability tests for the presence of features, not their organization. A fractured representation could contain every linearly decodable feature while being fractured in how those features relate to each other. The compositional parts are present but their composition is broken.

This connects directly to the "imposter intelligence" post angle: Can LLMs understand concepts they cannot apply?, Does supervised fine-tuning actually improve reasoning quality?, and Do foundation models learn world models or task-specific shortcuts?. All describe the same meta-pattern: surface metrics certify capability that internal structure analysis would disqualify.

The practical implication for model evaluation: passing compositional generalization tests does not guarantee robust compositional reasoning. Evaluation under distribution shift, perturbation, and novel recombination is required to distinguish genuine compositionality from fractured representations that happen to contain the right features.

Inquiring lines that read this note 185

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do stronger reasoning capabilities create tradeoffs with instruction following? Can inference-time compute effectively substitute for model scale? How effectively can language models perform reasoning, especially combined with symbolic methods? How do training data properties determine the emergence of internal misalignment? How do surface patterns enable correct outputs but reduce robustness? Can intelligent routing over smaller models outperform scaling a single large model? How do evaluation practices shape which failures stay visible? Can models improve accuracy without degrading reasoning quality? What causes reasoning models to fail or wander off track? Can mechanistic interpretability reliably guide practical model design choices? How do capability benchmark scores systematically misrepresent true model abilities? Why do embedding systems fail to capture task-relevant relationships? Do structural constraints outperform deep architectures in recommendation systems? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? Can compression size predict model complexity better than parameter count alone? Why does polished presentation create unearned authority in AI outputs? How should designers communicate what AI systems truly are and can do? What reasoning architectures enable models to solve complex problems efficiently? How does evaluation scope and dimensionality affect what we measure? What enables genuine semantic understanding in language models? How do neural networks achieve compositional generalization at scale? Do language models develop actual world models or merely task heuristics? What role does sparsity play in model behavior and scaling decisions? What compositional reasoning failures limit large language models despite scale? What capability trade-offs arise from domain specialization through fine-tuning? Do reasoning benchmarks predict model performance in long-horizon workflows? Can prompt-based context override biases that were embedded during pretraining? How does self-revision in reasoning models affect accuracy and confidence? How much do training data properties shape model reasoning? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Can reasoning scale in latent space without tokens? What training dynamics and scale trigger emergence of reasoning capabilities? Can harness architecture and protocols provide agent reliability without model scaling? Does alignment training create genuine alignment or just output compliance? Can brute-force automated research substitute for iterative depth and human research intuition? Why don't LLMs reliably translate capability into accurate outputs? How do prompting refinements mask underlying biases and model frequency patterns? How should items be represented and indexed in recommenders? Do language models respond to social pressure and face-saving like humans? What design and behavioral factors drive false consciousness attribution to AI? How does decomposing tasks improve reasoning and prevent failure propagation? How does persona conditioning amplify demographic stereotyping and bias in models? Where and how do personality traits reside in language models? Do reasoning traces faithfully reflect actual model reasoning? When do multi-agent systems outperform single frontier models? Do language models reason through causal mechanisms or semantic associations? When do semantic similarity approaches miss structural retrieval failures? What is the relationship between thinking tokens and reasoning accuracy? How can we prevent synthetic data from contaminating statistical inference and corpora? What training data selection strategies maximize generalization across difficulty levels? How do standardized protocols improve multi-agent coordination and reliability? How do pretraining biases affect reward signal effectiveness in RLVR? What do systematic disagreements between annotators reveal about ground truth? How does harness optimization generalize across different model architectures and domains? Is reasoning capability latent in base models or created by post-training? What makes distillation transfer some model capabilities while suppressing others? What structural properties of attention create systematic model biases? How should retrieval systems handle complex multi-step reasoning? Why do locally safe actions create system-level safety gaps? What attack surfaces do reasoning traces and chains introduce? Can we reliably detect when models game evaluations? Does encoded knowledge in language models actually influence their outputs? Does model confidence reliably signal actual accuracy in practice?

Related concepts in this collection 1

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 132 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

identical performance metrics can mask fundamentally different internal representations — feature linear decodability does not guarantee representational organization