Can models be smart without organized internal structure?
Explores whether linear feature decodability proves genuine compositional reasoning or merely indicates that the right features are present but poorly organized. Critical for understanding what performance metrics actually certify.
Two findings from mechanistic interpretability appear contradictory but operate at different levels of representational analysis:
Fractured Entangled Representations (FER): Since Can identical outputs hide broken internal representations?, SGD-trained models fail catastrophically under perturbation or distribution shift in ways that well-organized representations would not. The pathology is invisible to standard evaluation.
Compositional generalization at scale: Scaling data and model size produces representations where compositional features are linearly decodable — separable task constituents can be independently identified and manipulated. This has been taken as evidence for genuine compositional understanding.
The resolution: Linear decodability tests for the presence of features, not their organization. A fractured representation could contain every linearly decodable feature while being fractured in how those features relate to each other. The compositional parts are present but their composition is broken.
This connects directly to the "imposter intelligence" post angle: Can LLMs understand concepts they cannot apply?, Does supervised fine-tuning actually improve reasoning quality?, and Do foundation models learn world models or task-specific shortcuts?. All describe the same meta-pattern: surface metrics certify capability that internal structure analysis would disqualify.
The practical implication for model evaluation: passing compositional generalization tests does not guarantee robust compositional reasoning. Evaluation under distribution shift, perturbation, and novel recombination is required to distinguish genuine compositionality from fractured representations that happen to contain the right features.
Inquiring lines that read this note 185
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do stronger reasoning capabilities create tradeoffs with instruction following?- Why do only two of fourteen models improve when problem constraints are removed?
- How do unstated feasibility constraints affect model decision-making?
- Can a single SAE feature control reasoning behavior across model families?
- Why do models fail on logically equivalent tasks with different data distributions?
- Can you steer reasoning by directly manipulating SAE features?
- Can we steer model reasoning by manipulating single features?
- How does learnability at the observer's current state prevent novelty from breaking model reasoning?
- What structural constraints matter more than model depth for CF?
- How does fluent output mask the mythic function of a system?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- How should product specifications measure alignment without naming the dimension?
- What happens when alignment targets measure only the preferred dimension of entangled properties?
- Does parameter composition work when adapter alignment is imperfect?
- Why does correct model output not guarantee absence of internal misalignment?
- How do unstated constraints become invisible to training data distributions?
- Can bilevel autoresearch succeed when the inner and outer loops use different models?
- How much does domain shift limit the mechanisms a bilevel system can autonomously discover?
- When should model isolation be preferred over weight-averaging approaches?
- Why do power-law distributions make standard ML infrastructure assumptions fail?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Does model collapse occur across different architectures or only in specific conditions?
- How do overparameterization and data size shift what attractors represent?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- What production constraints should determine paradigm selection?
- Does model capability still matter once coordination infrastructure is optimized?
- What makes the frame problem distinct from feature-level shortcuts?
- How do autonomous pipelines identify and fix silent bugs in data pipelines?
- Why do unresolved items cluster in structured patterns rather than randomly?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- What happens when you remove core political features from a deep model?
- Why do models with less steerability have more abstract ideological features?
- Can mechanistic interpretability explain explanation-execution disconnection?
- How does mechanistic interpretability complement learning mechanics in explaining deep learning?
- What distinguishes a representational feature from a causally inert correlation?
- Can interventions on model components prove mechanism without explaining encoding?
- How do mechanistic features compare to natural language for interpretability?
- What makes representation engineering better than mechanistic interpretability for detecting hidden objectives?
- Why do feature visualizations alone fail to establish mechanistic claims?
- What makes some internal circuits more interpretable than others?
- Can mechanistic interpretability findings guide practical interventions in model design?
- How does representation engineering compare to mechanistic interpretability for auditing?
- How should benchmarks test whether models fit algorithms or patterns?
- How much do metric choices inflate claims about model capabilities?
- How do weight perturbations reveal what performance benchmarks cannot measure?
- Why do single function-calling benchmarks mask model weakness in specific areas?
- Why do text-only benchmarks underestimate deployed model capability?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How do coverage and identifiability set separate performance ceilings?
- Can a single Elo ranking represent multidimensional model capability?
- What capability dimensions does a single aggregate pass rate hide?
- How should single-axis benchmarks account for separable capability dimensions?
- Why do accuracy scores alone miss important dimensions of model capability?
- How do embedding dimension limits constrain what concept models can represent?
- How does discretization make item representations more distinguishable?
- Do multi-vector or cross-encoder models escape these dimensional constraints?
- Why is a combinatorial framework better than family resemblance classification?
- What architectural alternatives can capture compositional structure beyond pooled cosine?
- Can likelihood choice matter more than architectural depth for CF?
- What sparse high-rank patterns does the deep tower fail to capture?
- Why do cross-product features fail to generalize across unseen feature combinations?
- What compression explains why syntax fits in low-dimensional subspaces?
- Can steering vectors be combined with other compression techniques?
- How does requential coding measure true simplicity without parameter count inflation?
- Why do parameter-based compressors fail to measure true model simplicity?
- Can parameter compression mechanically force value systems toward idealized centers?
- Can Kolmogorov complexity alone capture what makes intelligence general?
- What makes AI-discovered architectures reveal design principles invisible to humans?
- How does nesting optimization levels improve on traditional network depth?
- What architectural properties of deterministic models block multi-solution reasoning?
- What role do multi-dimensional quality frameworks play in assessing arguments versus single-metric approaches?
- How does separating environment components make evaluation results more reproducible and analyzable?
- Is interpretive multiplicity a bug in language or a feature?
- How do functional features differ from representational abstract features?
- What test distinguishes genuine compositionality from fractured feature presence?
- What distinguishes conceptual understanding from statistical pattern matching in models?
- What distinct structural signatures do model repetition and topic volatility create?
- What spectral signatures distinguish hierarchy-driven geometry from corpus-driven geometry?
- Can we balance interpretability with the efficiency gains of compressed inter-model communication?
- Does architectural discovery follow an empirical scaling law like neural networks?
- Do larger models develop more abstract features than smaller ones?
- What makes linear decodability a reliable signal of compositionality?
- What makes a self-supervised pruning metric work without labels at scale?
- Can steering vectors prove that representations are genuinely organized?
- Can identical model performance mask fundamentally broken internal representations?
- Can fractured entangled representations hide undetected by standard analysis methods?
- Does the linear representation hypothesis reflect networks or reflect our analysis tools?
- Can representation engineering cleanly isolate single features in entangled semantic space?
- What are fractured entangled representations in neural networks?
- How do sparse circuits compare to the modular subnetworks that emerge naturally?
- Can geometric structure in representations exist without supporting functional mechanisms?
- Why does gradient descent discover compositional structure without explicit pressure?
- Do generic kernel-decay assumptions alone explain coarse-to-fine spectral ordering?
- Can spectral eigenvector ordering serve as a model-agnostic interpretability probe?
- Can representation analysis methods detect complex features models compute with?
- How can neural networks be interpretable by design rather than post-hoc?
- What physical structure does a Gaussian-regularized latent space actually encode?
- What makes regularization an implicit factor in embedding geometry?
- Do feature extraction methods systematically miss computationally important complex features?
- What makes a feature abstract versus concrete in neural network activations?
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- What prevents representation collapse in latent-prediction world models like JEPA?
- Can generative reconstruction preserve latent manifold structure better than geometric compression?
- Can a world model have rich representations without adequate data coverage?
- Why must world models be nested rather than flat and uniform?
- How do spectral-norm constraints prevent divergence in world model rollouts?
- What makes multimodal conditioning effective when features are decomposed to the right granularity?
- Why do singular value experts compose better than low-rank adapter subspaces?
- Why does weight sparsity reduce superposition and force disentangled representations?
- Can sparse approximations reveal interpretable structure hidden in existing dense models?
- Does sparsity enforce compositional structure or merely amplify existing modularity?
- Can sparsity patterns reliably indicate how well a model knows its input?
- How do sparse weight patterns affect model interpretability?
- What task structures benefit most from geometric parameter merging?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Why do metric choices constrain which model capabilities get developed?
- How can expensive models efficiently support cheap models in production?
- What distinctive properties make open foundation models different from closed ones?
- What benefits do open foundation models create that closed systems cannot?
- How much does workflow architecture matter versus raw model capability?
- What distinguishes new associations from existing ones at the computational level?
- What makes frozen model reasoning different from weight-based parameter updates?
- How does optimizing model performance decouple from optimizing user interpretability?
- Can a complexity-predictor be meaningful if models are redundant?
- Why does capturing domain structure reduce data requirements more than raw volume?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- What makes structured stochasticity more effective than unstructured randomness in reasoning?
- Why does the right structural prior matter more than raw model capacity?
- Why do energy-based models generalize better on out-of-distribution data than standard transformers?
- How does adjacent layer sharing differ from non-adjacent weight reuse?
- Why do standard transformers fail to encode recursive structure in their hidden states?
- Does Gemma's transformer explicitly exploit the inherited hierarchical geometry?
- How do pre-norm layers enable reliable fixed-point halting signals?
- How does LatentQA differ from predefined concept steering like representation engineering?
- What affordances do normalizing flows add over opaque vector reasoning?
- What skills can large models identify and organize about their own abilities?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- What makes some model capabilities reliable while others remain brittle?
- Why does increased model capability make detection harder in delegated workflows?
- Can end-to-end models maintain debuggability without modular components?
- Can alignment methods like DPO exploit or correct these surface feature biases?
- Does token-level loss aggregation help aligned models differently?
- Can structured decomposition fix evaluation gaps in other research tasks?
- Can surface-level correctness hide failures in structural learning by LLMs?
- Can granular function calling tasks learn composition from graph-sampled data?
- Can we predict which tasks will decompose into modular subnetworks?
- When does backward decomposition fail on open-ended or unstructured tasks?
- What makes well-formatted outputs misleading as evidence of model capability?
- Why do semi-formal templates improve verification accuracy over unstructured reasoning?
- Can entropy signatures alone detect whether context was model-generated or externally prefilled?
- Can seedless generation maintain explainability while scaling control?
Related concepts in this collection 1
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can we track and steer personality shifts during model finetuning?
This research explores whether personality traits in language models occupy specific linear directions in activation space, and whether we can detect and control unwanted personality changes during training using these geometric directions.
persona vectors demonstrate a case where linear decodability corresponds to genuine functional organization (steering works), providing a positive counterexample to FER's warning that decodability alone is insufficient
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Titans: Learning to Memorize at Test Time
- Break It Down: Evidence for Structural Compositionality in Neural Networks
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- Large Language Model Reasoning Failures
Original note title
identical performance metrics can mask fundamentally different internal representations — feature linear decodability does not guarantee representational organization