Nested Learning: The Illusion of Deep Learning Architecture Expanded
In the previous sections, we discussed the concept of nested learning and how existing well-known components of neural networks such as popular optimizers and architectures fall under the NL paradigm. In this section, we discuss the takeaways, connection of different concepts, and the implications of NL perspective on common terms.
Memory and Learning. For a long period of time, in machine learning models, memory have been treated as a separate block with a clear distinction between its parameters and other components. Such designs often assume a short and/or long-term memory blocks, where short-term memory is responsible for the local context, while long-term memory is the storage for the persistent knowledge in models. In human brain, however, memory is considered as a distributed interconnected system without a clear known components that are independently responsible for short or long-term memory. In NL, we build upon a common terminology for memory and learning in neuropsychology literature, indicating that: Memory is a neural update caused by an input and so learning is the process of acquiring useful memory (Okano et al. 2000). From this viewpoint, any update by gradient descent (or any other optimization algorithms) in any levels of neural learning module is considered as a form of memory. Interestingly, our findings in Section 4.1 on gradient descent being (self-referential) associative memory is aligned with this terminology. Furthermore, based on this terminology, in continuum memory system, the neural updates are applied at different frequencies and so memories are stored with different time scales, resulting in more robust memory management with respect to catastrophic forgetting.
Memory and Learning from Nested Learning Perspective: Memory is not an isolated system and is distributed throughout the parameters. Particularly, any update that is caused by the input is a stored memory in the neural network, and the process of effectively store, encode, and in general acquire such memories is referred to as learning process.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How should agents manage memory granularity to improve long-term performance?- How should future memory systems control what gets written and trusted?
- How do the three-axis taxonomies of memory forms and functions differ?
- When does training a memory model beat RAG or fine-tuning?
- Can a memory module be swapped between different base models?
- Can native memory procedures acquired through training handle stale or incorrect cached information?
- Why do accumulated memory systems sometimes hurt continual learning?
- Why do external memory consolidation systems fail worse than naive in-context learning on continual tasks?
- What counts as genuine memory under the Extended Mind thesis?
- Why do CoALA and Letta disagree on what counts as working memory?
- How does co-activation shape which memories become linked together?
- Why does LLM memory consolidation regress below no-memory baselines?
- How does consolidation schedule order affect final memory quality?
- How does the hippocampus bind disparate elements without storing everything itself?
- How does continuous implicit memory formation differ from explicit memory encoding?
- Does composing multiple continual learning mechanisms reduce forgetting more than single approaches?
- Can continuum memory systems prevent catastrophic forgetting in neural networks?
- How does dual-rate learning separate episodic and procedural memory in neural networks?
- How do cortical columns implement local inference over memory cycles?
- How do memorization and attention map onto different memory systems?
- Why do hybrid memory systems outperform single-tier AI architectures?