SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Can proximity between teacher and student fix distillation instability?

On-policy distillation works well in theory but fails in practice due to capacity gaps. Does dynamically constructing a proximal teacher within a trust region resolve this fragility?

Synthesis note · 2026-07-17 · sourced from Reinforcement Learning

On-policy distillation (OPD) has become the default LLM post-training paradigm because it occupies a sweet spot: it is mathematically equivalent to RL where the immediate reward is the log-probability ratio between teacher and student policies, so its on-policy nature avoids the catastrophic forgetting of SFT while its dense rewards escape the sample inefficiency and instability of RLVR (which gives only a sparse verifiable signal at the end). But standard OPD is optimization-fragile in practice, and TOP-D locates the bottleneck precisely: the capacity gap between a strong target teacher and a weaker student produces high-variance, unstable gradients.

The fix is to stop distilling directly from the distant target teacher and instead dynamically construct a proximal teacher — a teacher close to the current student — and iterate within a trust region. Theoretically this inherently controls gradient variance and, with safe off-policy data reuse inside the trust-region iterations, yields a formal global-convergence result and a monotonic-improvement bound. The framing is deliberately the classic RL move (TRPO's trust region) transplanted onto distillation: bound how far each update moves the policy, and stability follows for free — TOP-D adds zero computational overhead, decisively outperforming standard OPD and competitive RLVR baselines on mathematical reasoning across scales.

This connects to a growing recognition that the teacher/student gap is the load-bearing variable in distillation, not raw teacher strength. Since Does richer teacher context hurt student generalization?, the intuition that a maximally strong or maximally informed teacher is best keeps failing; proximity, not power, governs whether the transferred signal is learnable.

Inquiring lines that read this note 15

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What training data selection strategies maximize generalization across difficulty levels? How do surface patterns enable correct outputs but reduce robustness? What makes distillation transfer some model capabilities while suppressing others? How does harness optimization generalize across different model architectures and domains? Can we reliably detect when models game evaluations? How do pretraining biases affect reward signal effectiveness in RLVR? What capability trade-offs arise from domain specialization through fine-tuning? How do agent-learned skills transfer and improve across different tasks?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 103 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

a dynamically constructed proximal teacher with a trust region turns unstable on-policy distillation into a stable monotonically improving paradigm