Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Paper · Source
Mechanistic Interpretability

Mechanistic interpretability seeks to understand neural networks by breaking them into components that are more easily understood than the whole. By understanding the function of each component, and how they interact, we hope to be able to reason about the behavior of the entire network. The first step in that program is to identify the correct components to analyze.

Unfortunately, the most natural computational unit of the neural network – the neuron itself – turns out not to be a natural unit for human understanding. This is because many neurons are polysemantic: they respond to mixtures of seemingly unrelated inputs. In the vision model Inception v1, a single neuron responds to faces of cats and fronts of cars [1] . In a small language model we discuss in this paper, a single neuron responds to a mixture of academic citations, English dialogue, HTTP requests, and Korean text. Polysemanticity makes it difficult to reason about the behavior of the network in terms of the activity of individual neurons.

One potential cause of polysemanticity is superposition [2, 3, 4, 5] , a hypothesized phenomenon where a neural network represents more independent "features" of the data than it has neurons by assigning each feature its own linear combination of neurons. If we view each feature as a vector over the neurons, then the set of features form an overcomplete linear basis for the activations of the network neurons. In our previous paper on Toy Models of Superposition [5] , we showed that superposition can arise naturally during the course of neural network training if the set of features useful to a model are sparse in the training data. As in compressed sensing, sparsity allows a model to disambiguate which combination of features produced any given activation vector. 1

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does decomposing tasks improve reasoning and prevent failure propagation? What enables genuine semantic understanding in language models? Do language models develop actual world models or merely task heuristics? Can memory architectures handle ultra-long context better than attention? What training dynamics and scale trigger emergence of reasoning capabilities? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How do neural networks achieve compositional generalization at scale? What reasoning architectures enable models to solve complex problems efficiently? What role does sparsity play in model behavior and scaling decisions? What compositional reasoning failures limit large language models despite scale? How should designers communicate what AI systems truly are and can do? Can mechanistic interpretability reliably guide practical model design choices?