SYNTHESIS NOTE
Topics›LLM Architecture›this note

Can neural memory modules scale language models beyond attention limits?

Can separating short-term attention from adaptive long-term memory allow models to efficiently handle context windows exceeding 2M tokens while maintaining competitive performance?

Synthesis note · 2026-02-22 · sourced from LLM Architecture

Titans (2501.00663) introduces a neural long-term memory module that addresses a fundamental contradiction in linear recurrent models: they are designed for efficiency on long contexts, but long contexts cannot be properly compressed into small fixed-size states.

The architectural insight is that attention and memory serve fundamentally different functions. Attention operates as short-term memory — accurate direct dependency modeling within the current context window, but quadratic cost limits its reach. Neural memory operates as long-term memory — compressed and persistent, memorizing data that is surprising or close to surprising tokens. The memory update mechanism considers the proportion of memory size to data surprise, resulting in adaptive memory management.

Three integration variants are proposed: memory as context (attending to memory alongside current context), memory as gating (memory modulates attention output), and memory as a layer (memory replaces some attention layers). Each variant trades off between integration depth and computational overhead.

The results establish that Titans outperform both standard Transformers (with the same context window) and modern linear recurrent models across language modeling, common-sense reasoning, genomics, and time series. Critically, Titans scale to context windows larger than 2M tokens while showing competitive performance with Transformers that use the entire context — the long-context problem is addressed without the quadratic penalty. The persistent nature of the memory module makes it a natural substrate for Can models precompute answers before users ask questions? — the memory can store precomputed inferences between interactions, and sleep-time processing can populate the memory with anticipated query-relevant information.

Since Can models reason without generating visible thinking tokens?, the Titans architecture offers a complementary path: rather than scaling reasoning depth through recurrent computation, it scales memory breadth through adaptive memorization. Both bypass the limitations of standard attention but along different architectural dimensions.

Miras unifying framework (2504.13173): The "It's All Connected" paper reconceptualizes Transformers, Titans, and modern linear recurrent models as associative memory modules that learn a mapping of keys to values using an internal objective — termed "attentional bias." The paper observes that most existing sequence models use either dot-product similarity or ℓ2 regression objectives as their attentional bias. Miras provides a general framework with four design choices: (i) associative memory architecture, (ii) attentional bias objective, (iii) retention gate, and (iv) memory learning algorithm. Forgetting mechanisms are reinterpreted as retention regularization — providing a principled basis for forget gates across architectures. Three novel sequence models — Moneta, Yaad, and Memora — go beyond existing linear RNNs while maintaining fast parallelizable training. Different Miras configurations yield models with varying strengths: some excel at language modeling, others at commonsense reasoning or recall-intensive tasks. This generalizes the Titans insight: the attention-as-short-term/memory-as-long-term distinction is one instance of a broader design space where attentional bias objective and retention mechanism can be independently varied.

Inquiring lines that read this note 145

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What design and behavioral factors drive false consciousness attribution to AI? What prevents conversational agents from taking initiative in dialogue? What structural properties of attention create systematic model biases? What compositional reasoning failures limit large language models despite scale? Can compression size predict model complexity better than parameter count alone? Can memory architectures handle ultra-long context better than attention? Do language models learn genuine understanding or just surface patterns? Can diffusion models match autoregressive performance on language generation tasks? What role does sparsity play in model behavior and scaling decisions? Why does adding new knowledge through fine-tuning degrade existing capabilities? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? What determines appropriate intervention timing and manner for AI agents? How should retrieval systems handle complex multi-step reasoning? How do surface patterns enable correct outputs but reduce robustness? How should inference compute be allocated based on problem difficulty? Can prompt-based context override biases that were embedded during pretraining? Why do token-level mechanisms matter for learning to reason? How should systems decide whether to retrieve or reason alone? Do structural constraints outperform deep architectures in recommendation systems? Do reasoning benchmarks predict model performance in long-horizon workflows? Why does memory consolidation cause performance regression in continual learning? How do neural networks achieve compositional generalization at scale? What causes reasoning models to fail or wander off track? Can inference-time compute effectively substitute for model scale? How much do training data properties shape model reasoning? Does abstract user knowledge outperform concrete interaction history in personalization? Why can't prompting alone inject genuinely new knowledge into models? Does RL create genuinely new reasoning capabilities or refine existing ones? Can reasoning scale in latent space without tokens? How should agents manage memory granularity to improve long-term performance? How do agent-learned skills transfer and improve across different tasks?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 116 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

neural memory modules that adaptively memorize surprising tokens complement attention as long-term vs short-term memory — scaling to 2M+ context