SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Can decoding-time tuning preserve knowledge better than weight fine-tuning?

Explores whether applying alignment signals at inference time rather than modifying model weights can better preserve the factual knowledge learned during pretraining while still achieving alignment goals.

Synthesis note · 2026-02-22 · sourced from Training Fine Tuning

Proxy-tuning fine-tunes a small model, then applies the difference between the small tuned and small untuned model's predictions to shift a large untuned model's outputs at decoding time. The large model's parameters are never modified. The method closes 91% of the performance gap between Llama-2-13B and its directly tuned CHAT version, and 88% for the 70B model.

The critical finding: on knowledge-intensive tasks, proxy-tuning sometimes surpasses the performance of direct instruction-tuning. This is because direct fine-tuning modifies model weights — and some of those modifications overwrite pretrained knowledge. Since Why does reasoning training help math but hurt medical tasks?, weight modification risks corrupting the knowledge storage that proxy-tuning leaves intact.

Proxy-tuning primarily promotes reasoning and stylistic tokens. Analysis of the token-level distributional shift shows the largest influence on tokens associated with reasoning patterns and output style — consistent with evidence that "alignment mainly affects style rather than knowledge." This aligns with Does instruction tuning teach task understanding or output format? and Can imitating ChatGPT fool evaluators into thinking models improved?: what fine-tuning actually changes is output distribution, not capability. Proxy-tuning achieves this distributional change without touching the model weights that encode knowledge.

For domain adaptation, proxy-tuning Llama-2-13B using CodeLlama-7B produces 17-32% improvement on coding benchmarks. The small expert provides the distributional guidance; the large base model provides the knowledge. An optional hyperparameter controls the amount of guidance, enabling runtime trade-offs between different generation attributes.

This constitutes a fifth paradigm in the How do knowledge injection methods trade off flexibility and cost?: decoding-time adaptation. Zero training cost on the target model, full knowledge preservation, but requires access to base model logits at inference time.

ARGS (Alignment as Reward-Guided Search) provides a complementary inference-time method. Instead of applying a distributional shift from a tuned proxy, ARGS adjusts model predictions at each decoding step using a reward signal directly. Two components: reward-guided scoring (assigns scores to possible continuations) and token selection (selects a continuation based on scored candidates). A tunable weight controls the trade-off between semantic relevance and alignment criteria — setting it to zero recovers standard maximum-likelihood decoding. ARGS enables rapid personalized alignment without retraining: different users can have different reward functions applied at inference time. Together, proxy-tuning (distributional shift from expert delta) and ARGS (reward-guided decoding) suggest a design space where multiple axes of adaptation — domain knowledge, user preferences, task constraints — can each be applied at decoding time through complementary mechanisms. See Can user preferences be learned from just ten questions? for how per-user reward functions can be efficiently constructed.

Inquiring lines that read this note 156

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does alignment training create genuine alignment or just output compliance? Can self-generated feedback reliably guide model training without ground truth? Can local safety checks guarantee system-level behavioral safety? How do training data properties determine the emergence of internal misalignment? Can memory architectures handle ultra-long context better than attention? Why does adding new knowledge through fine-tuning degrade existing capabilities? How do pretraining biases affect reward signal effectiveness in RLVR? What capability trade-offs arise from domain specialization through fine-tuning? Why do stronger reasoning capabilities create tradeoffs with instruction following? Can models improve accuracy without degrading reasoning quality? What emerges when safety-aligned models attempt to role-play deceptive personas? What training dynamics and scale trigger emergence of reasoning capabilities? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What training data selection strategies maximize generalization across difficulty levels? What role does sparsity play in model behavior and scaling decisions? What makes distillation transfer some model capabilities while suppressing others? What compositional reasoning failures limit large language models despite scale? How do surface patterns enable correct outputs but reduce robustness? Can inference-time compute effectively substitute for model scale? Can prompt-based context override biases that were embedded during pretraining? Can inoculation prompting prevent emergent misalignment after reward hacking? Should agents decouple planning from perception grounding for better performance? Why do token-level mechanisms matter for learning to reason? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can iterative DPO replicate online reinforcement learning dynamics for research? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? How should inference compute be allocated based on problem difficulty? Do language models learn genuine understanding or just surface patterns? How much do training data properties shape model reasoning? How do prompting refinements mask underlying biases and model frequency patterns? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can compression size predict model complexity better than parameter count alone? How do neural networks achieve compositional generalization at scale? How does harness optimization generalize across different model architectures and domains? Can intelligent routing over smaller models outperform scaling a single large model? Can mechanistic interpretability reliably guide practical model design choices? Why do embedding systems fail to capture task-relevant relationships? Why does memory consolidation cause performance regression in continual learning? Why can't prompting alone inject genuinely new knowledge into models? Why don't LLMs reliably translate capability into accurate outputs? Does RL create genuinely new reasoning capabilities or refine existing ones? Does preference optimization systematically degrade conversational grounding in language models? How does synthetic data quality and diversity affect downstream model capabilities? How should systems decide whether to retrieve or reason alone? What enables genuine semantic understanding in language models? How do agent-learned skills transfer and improve across different tasks?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 186 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

proxy tuning at decoding time preserves pretrained knowledge better than direct fine-tuning by applying the tuning signal as a distributional shift