SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Does instruction tuning teach task understanding or output format?

Exploring whether models trained on instructions actually learn the task semantics or merely learn to match output distributions. This matters because it challenges assumptions about how fine-tuning improves model behavior.

Synthesis note · 2026-02-22 · sourced from Training Fine Tuning

"Do Models Really Learn to Follow Instructions?" creates two devastating controls. First, simplified task definitions that strip all semantic content, leaving only output space information (e.g., "output one of: A, B, C"). Second, delusive examples containing incorrect input-output mappings. Models trained on either achieve comparable performance to models trained on full, correct instructions. A random baseline achieves 42.6% exact-match versus instruction tuning's 43%.

The implication: instruction tuning primarily teaches the model to map its existing capabilities to the expected output format, not to understand or execute the task as described in the instruction. The semantic content of the instruction — what the task is, how to approach it, what constitutes a correct answer — appears largely irrelevant. What matters is the output distribution: how many classes, what format, what vocabulary.

This connects to a broader pattern. Does training data format shape reasoning strategy more than domain? showed a 7.5x stronger effect of format over domain. Can models pass tests while missing the actual grammar? showed that correct outputs can mask reliance on surface heuristics. The instruction tuning finding adds: even explicit instructions about the task are largely ignored in favor of format signals.

A complementary theory from "Are Emergent Abilities just ICL?" (2309.01809) provides the mechanistic explanation: instruction tuning enables "implicit in-context learning" — mapping instructions to the form required for ICL rather than creating new functional abilities. The evidence: purported emergent abilities are explained by a combination of in-context learning, model memory, and linguistic knowledge. The model's sensitivity to minor prompt variations and tendency to hallucinate are inconsistent with genuine emergent functional abilities but consistent with a model that maps prompts to ICL patterns. This reframes safety concerns: if prompts function as "training mechanisms" rather than interfaces to inherent abilities, the safety landscape changes — the risk is in what ICL patterns exist, not in what abilities have "emerged."

The IT Survey (same source) documents the concern from the other direction: "there has been an intense criticism that IT only captures surface-level patterns and styles rather than comprehending and learning the task." Combined with the False Promise finding that model imitation captures style not factuality, a clear pattern emerges: fine-tuning-based adaptation — whether through imitation, instruction tuning, or domain SFT — preferentially captures distributional and formatting information while leaving underlying capabilities largely unchanged. The capability bottleneck is in the base model, not the adaptation method.

Webson & Pavlick (2021) provide the prompting-level parallel. Evaluating 30+ manually written templates and 13 sets of target words across 390+ prompts, they find models learn identically fast from irrelevant or misleading templates as from instructive ones. Models are "much more sensitive to the choice of LM target words as opposed to the meaning of the instruction templates." Instruction-tuned models can be "too robust" — less sensitive to prompt semantics than non-IT equivalents, suggesting IT trains a form of prompt-blindness. This holds from 235M to 175B parameters. The convergence is striking: both the fine-tuning and the prompting literature arrive at the same conclusion from opposite directions — the semantic content of instructions is largely inert, and what transfers is format and output space information.

Inquiring lines that read this note 167

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance promote real skill development or substitute for independent learning? What training dynamics and scale trigger emergence of reasoning capabilities? How does decomposing tasks improve reasoning and prevent failure propagation? Do reasoning benchmarks predict model performance in long-horizon workflows? Why do stronger reasoning capabilities create tradeoffs with instruction following? Does encoded knowledge in language models actually influence their outputs? What training data selection strategies maximize generalization across difficulty levels? Does alignment training create genuine alignment or just output compliance? How do training data properties determine the emergence of internal misalignment? Do language models develop actual world models or merely task heuristics? Why can't prompting alone inject genuinely new knowledge into models? How much do training data properties shape model reasoning? How does AI adoption across firms reshape employment and inequality? What determines appropriate intervention timing and manner for AI agents? Can prompt-based context override biases that were embedded during pretraining? Is reasoning capability latent in base models or created by post-training? How do prompting refinements mask underlying biases and model frequency patterns? Does RL create genuinely new reasoning capabilities or refine existing ones? Can models improve accuracy without degrading reasoning quality? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Do language models learn genuine understanding or just surface patterns? How can we prevent synthetic data from contaminating statistical inference and corpora? How effectively can language models perform reasoning, especially combined with symbolic methods? What capability trade-offs arise from domain specialization through fine-tuning? Why do persona simulations fail to predict authentic user behavior? Can iterative DPO replicate online reinforcement learning dynamics for research? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Can self-generated feedback reliably guide model training without ground truth? What makes step-level supervision effective for complex reasoning traces? Should GUI agents use structured representations over raw visual input? Should agents decouple planning from perception grounding for better performance? What role does sparsity play in model behavior and scaling decisions? How do spurious versus genuine rewards shape model reasoning and behavior? How do pretraining biases affect reward signal effectiveness in RLVR? Why does adding new knowledge through fine-tuning degrade existing capabilities? Can inoculation prompting prevent emergent misalignment after reward hacking? How do prompt design choices influence model reasoning and performance? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Can mechanistic interpretability reliably guide practical model design choices? Why do embedding systems fail to capture task-relevant relationships? Why don't LLMs reliably translate capability into accurate outputs? How do capability benchmark scores systematically misrepresent true model abilities? Why do some clarifying approaches produce understanding while others just satisfy? What attack surfaces do reasoning traces and chains introduce? How do agent-learned skills transfer and improve across different tasks? How does harness optimization generalize across different model architectures and domains? How do neural networks achieve compositional generalization at scale? Can brute-force automated research substitute for iterative depth and human research intuition? Can reasoning scale in latent space without tokens? How does reasoning length affect model performance across different tasks? How should test-time compute scaling work in agentic systems? Can intelligent routing over smaller models outperform scaling a single large model? Does model confidence reliably signal actual accuracy in practice? Do reasoning traces faithfully reflect actual model reasoning?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 168 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

instruction tuning teaches output format distribution not task understanding — simplified and delusive instructions achieve comparable performance