SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Does staying close to the base model preserve learning ability?

Explores whether limiting how far training pushes a model from its base distribution (measured by KL divergence) helps it learn new tasks more effectively over time, and why that trade-off matters for continual learning.

Synthesis note · 2026-05-28 · sourced from Training Fine Tuning

There is a quiet variable connecting forgetting, generalization, and the ability to keep learning: how far training pushes the policy from its base distribution, measured as KL divergence. The Fast-Slow result makes the relationship explicit. FST-trained models stay up to 70% closer to the base LLM in KL than parameter-only RL — and that reduced drift is not just a forgetting story. It preserves plasticity: after training on one task, FST models adapt more effectively to a subsequent task, while parameter-only RL stalls when task domains change on the fly.

The pattern is that drift and plasticity trade off. Each parameter update that improves in-domain reward also moves the model toward a sharper, lower-entropy policy specialized to that task. Specialization is exactly what makes the model less able to absorb the next task — the weights have committed. By keeping most task-specific adaptation in the fast textual channel and letting the slow weights move only a little, FST holds the policy near its flexible base, where it retains the entropy and breadth needed to learn again. Low KL drift is the leading indicator; preserved plasticity and reduced forgetting are downstream consequences.

Why it matters: it gives continual learning a measurable target. Rather than treating "don't forget" and "stay adaptable" as separate desiderata to engineer, you can watch a single quantity — distance from base — and recognize that overshooting it is what produces both forgetting and plasticity loss. It also reframes KL regularization (already standard in RLHF as a leash) as not merely a stability or alignment-preservation device but as the mechanism that keeps the model trainable in the future. The counterpoint: staying near base also caps how much any single task can specialize the weights, so for a one-shot deployment with no future tasks, aggressive drift may be the better trade.

Inquiring lines that read this note 73

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What capability trade-offs arise from domain specialization through fine-tuning? Why does adding new knowledge through fine-tuning degrade existing capabilities? Do language models develop actual world models or merely task heuristics? What training dynamics and scale trigger emergence of reasoning capabilities? What makes distillation transfer some model capabilities while suppressing others? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What training data selection strategies maximize generalization across difficulty levels? How do surface patterns enable correct outputs but reduce robustness? How much do training data properties shape model reasoning? Does RL create genuinely new reasoning capabilities or refine existing ones? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can prompt-based context override biases that were embedded during pretraining? Can compression size predict model complexity better than parameter count alone? How do neural networks achieve compositional generalization at scale? What role does sparsity play in model behavior and scaling decisions? How do pretraining biases affect reward signal effectiveness in RLVR? Does preference optimization systematically degrade conversational grounding in language models? Can memory architectures handle ultra-long context better than attention? Why does memory consolidation cause performance regression in continual learning? How do agent-learned skills transfer and improve across different tasks? What factors drive AI persuasiveness and how can it be mitigated? Does AI assistance promote real skill development or substitute for independent learning? Is reasoning capability latent in base models or created by post-training? Why do stronger reasoning capabilities create tradeoffs with instruction following? How does harness optimization generalize across different model architectures and domains?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 117 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

lower kl drift from the base model preserves plasticity enabling stronger continual learning on later tasks