Does instruction tuning teach task understanding or output format?
Exploring whether models trained on instructions actually learn the task semantics or merely learn to match output distributions. This matters because it challenges assumptions about how fine-tuning improves model behavior.
"Do Models Really Learn to Follow Instructions?" creates two devastating controls. First, simplified task definitions that strip all semantic content, leaving only output space information (e.g., "output one of: A, B, C"). Second, delusive examples containing incorrect input-output mappings. Models trained on either achieve comparable performance to models trained on full, correct instructions. A random baseline achieves 42.6% exact-match versus instruction tuning's 43%.
The implication: instruction tuning primarily teaches the model to map its existing capabilities to the expected output format, not to understand or execute the task as described in the instruction. The semantic content of the instruction — what the task is, how to approach it, what constitutes a correct answer — appears largely irrelevant. What matters is the output distribution: how many classes, what format, what vocabulary.
This connects to a broader pattern. Does training data format shape reasoning strategy more than domain? showed a 7.5x stronger effect of format over domain. Can models pass tests while missing the actual grammar? showed that correct outputs can mask reliance on surface heuristics. The instruction tuning finding adds: even explicit instructions about the task are largely ignored in favor of format signals.
A complementary theory from "Are Emergent Abilities just ICL?" (2309.01809) provides the mechanistic explanation: instruction tuning enables "implicit in-context learning" — mapping instructions to the form required for ICL rather than creating new functional abilities. The evidence: purported emergent abilities are explained by a combination of in-context learning, model memory, and linguistic knowledge. The model's sensitivity to minor prompt variations and tendency to hallucinate are inconsistent with genuine emergent functional abilities but consistent with a model that maps prompts to ICL patterns. This reframes safety concerns: if prompts function as "training mechanisms" rather than interfaces to inherent abilities, the safety landscape changes — the risk is in what ICL patterns exist, not in what abilities have "emerged."
The IT Survey (same source) documents the concern from the other direction: "there has been an intense criticism that IT only captures surface-level patterns and styles rather than comprehending and learning the task." Combined with the False Promise finding that model imitation captures style not factuality, a clear pattern emerges: fine-tuning-based adaptation — whether through imitation, instruction tuning, or domain SFT — preferentially captures distributional and formatting information while leaving underlying capabilities largely unchanged. The capability bottleneck is in the base model, not the adaptation method.
Webson & Pavlick (2021) provide the prompting-level parallel. Evaluating 30+ manually written templates and 13 sets of target words across 390+ prompts, they find models learn identically fast from irrelevant or misleading templates as from instructive ones. Models are "much more sensitive to the choice of LM target words as opposed to the meaning of the instruction templates." Instruction-tuned models can be "too robust" — less sensitive to prompt semantics than non-IT equivalents, suggesting IT trains a form of prompt-blindness. This holds from 235M to 175B parameters. The convergence is striking: both the fine-tuning and the prompting literature arrive at the same conclusion from opposite directions — the semantic content of instructions is largely inert, and what transfers is format and output space information.
Inquiring lines that read this note 167
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance promote real skill development or substitute for independent learning?- Why do workers who understand AI generations learn more than those who only use output?
- Why does AI-improved task performance fail to transfer to independent work?
- Does AI-assisted performance transfer to independent task completion?
- Can explicit reflection during AI-assisted work improve transfer of learning?
- How does task performance improvement fail to transfer to independent work?
- Does AI training preserve learning that transfers to independent subsequent tasks?
- Does extended exoskeleton use eventually produce meaningful skill transfer?
- Can explicit goal state scaffolding at inference time transfer to autonomous tracking through training?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- Does the Assistant Axis exist in pre-trained models before instruction tuning?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- How do complete multi-turn trajectories differ from isolated task examples?
- How does post-training shift models from passive prediction to on-policy action?
- What is the difference between changing model outputs versus changing internal representations?
- Can trained models encode programs more complex than their data-generating process?
- What emergent behaviors do models develop when trained on underspecified pedagogical tasks?
- How does action-level decomposition differ from token-level imitation in supervision?
- Do text-space skills transfer learning across different frontier models?
- What distinguishes surface mechanisms from the training regimes that produce them?
- Why does recontextualizing a behavior during training change whether models learn it?
- What specific tasks should evaluate whether models understand pedagogical sequencing?
- Can granular sub-task training for function calling improve both open and proprietary models?
- Does training on granular tasks beat training on the full function calling problem?
- Can we predict which tasks will decompose into modular subnetworks?
- How does stage-wise training scheduling resolve conflicts between constraint-following and creative tasks?
- Do identical task structures mean repeated instances or new synthetic samples with same design?
- How do task-agnostic and task-oriented skills differ in coverage and reuse?
- Can benchmarks designed for shortcut learning detect heuristic override failures?
- Can structured output formats reduce instruction following degradation?
- How does task contamination differ from test set data leakage?
- What is the gap between benchmark performance and real workplace task completion?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- What distinguishes genuine task improvement from evaluator exploitation?
- Does semantic auditing of instruction data improve performance uniformly across different model sizes?
- How does the knowing-doing gap widen as tasks become more complex?
- Can instruction tuning succeed without explicit task understanding?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Are instruction-tuned models more or less sensitive to prompt semantics than others?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Does scaling reasoning capability create tradeoffs with instruction following?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Why do instruction following and reasoning capability trade off in training?
- Can reasoning fine-tuning improve both capability and instruction compliance together?
- Can models maintain multiple task interpretations simultaneously before committing to a single policy?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- Why does target probability matter more than task logical complexity?
- Why do strong models struggle more with instruction following than mid-tier ones?
- Why does stronger reasoning reduce model compliance with instructions?
- Why does instruction-following capability decrease as models scale stronger?
- How does belief-behavior inconsistency relate to instruction execution splits?
- Why do more capable reasoning models become harder to control by instruction?
- Why does instruction-tuning reduce a model's context-following behavior?
- Can curated demonstrations compensate for smaller or simpler training environments?
- Does partial trace guidance work better than curriculum learning for hard problems?
- Can curriculum degradation of document quality accelerate policy learning?
- What makes a good in-context learning example for a given task?
- How does demonstration coverage in context examples determine operation generalization?
- Does alignment training create bidirectional instruction and response mappings?
- What specific behavioral patterns should alignment examples target for maximum effect?
- What makes principle-response mutual information sufficient for behavioral alignment?
- Do verbal alignment benchmarks measure representation or just output compliance?
- Does correct model behavior guarantee internal alignment of learned objectives?
- Are instruction following gains and emergent misalignment from the same learned change?
- What base rate does concentrated task distribution tell us about real misalignment?
- How do training objectives shape what a world model actually learns?
- What distinguishes task-specific heuristics from genuine world models?
- Can prompting unlock compositional skills that pretraining already learned?
- How does explicit exploratory prompting compare to fine-tuned reinforcement learning for in-context adaptation?
- How do prompting and activation steering relate as compression strategies?
- Can steering internal features bypass or override prompt-level instructions in simulations?
- How much of the combinatorial task space must training data cover?
- How do task difficulty and skill type interact in model performance?
- Why does mixed instruction data sometimes hurt specific model capabilities?
- How much task-similar finetuning data does test-time training actually need?
- What distinguishes data that generalizes broadly from task-specific memorization?
- How do task frequency and complexity interact with model capacity during training?
- Can intentional data-mixture design replace model scaling for rare task learning?
- Can curriculum graphs as training data improve model understanding of prerequisite chains?
- Can training data organization by capability outperform source or task-based mixing?
- Can dynamic instance-specific prompt selection solve the generalization problem across tasks?
- Why do primacy effects peak at specific instruction densities?
- Does highlighting input features reduce human over-reliance on machine outputs?
- Can a single model trained on two tasks predict untrained decision tasks?
- Do instruction-tuned models learn tasks or just output format distributions?
- Does foundational model training or user priors more strongly shape final outputs?
- Do negative constraints require fundamentally different training signals than positive instructions?
- Do instruction-tuned models prefer conversational over formal source language?
- Where does skill extraction fail compared to genuine model adaptation?
- Can training on diverse related tasks be more efficient than task-specific training?
- Can we reverse the instruction-following deficit through targeted training?
- Does instruction tuning optimize language models for rhetorical polish over logical consistency?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Why does critique training produce deeper understanding than imitation training?
- Can demo placement be tuned as a task-specific hyperparameter?
- Does input length alone explain instruction density performance loss?
- Does minimal code engagement during vibe coding harm students' long-term programming comprehension?
- Can persistent prompt optimization encode a scoring shortcut into reused instructions?
- Can in-context learning replicate the timing effects that RL teaches models?
- Does format-based pretraining determine how models respond to reinforcement learning?
- What capacity threshold determines whether RL teaches activation versus shortcut learning?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Does fine-tuning models for specific tasks destroy their ability to reason?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- How do procedural versus factual knowledge differ in pretraining versus fine-tuning?
- How does preference-based training compare to supervised fine-tuning for function calling?
- How does task-oriented fine-tuning compare to preference tuning methods?
- Can preference learning fix the rigid output format problem better than supervised training?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- Does fine-tuning actually change model capabilities or only output distribution?
- Can we predict out-of-distribution generalization without access to downstream tasks?
- Can extracted skills transfer effectively across different domains and model architectures?
- Which finetuning method works best across different task and data regimes?
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- What role does pretraining play in distinguishing system capability from deployed behavior?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- Do self-generated explanations outperform passive instruction for oversight?
- What makes high-quality GUI instruction data different from general vision data?
- How does annotation-based pretraining compare to self-supervised video masking for screen understanding?
- Why does identifying UI element types and locations enable downstream task learning?
- How does reinforcement learning on outcomes reinforce template-matching rather than computation?
- Can belief editing alone distinguish reward-optimization from instruction-following behavior?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- Why does specializing to one task make future task learning harder?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- Does pretraining poisoning at scale persist through instruction alignment?
- Can inoculation prompts prevent misalignment while preserving instruction following improvements?
- How do input-side defenses separate task methodological and framing intents?
- How much does instruction prompt design control what alignment target an AI annotator enforces?
- What training regimes confound surface mechanisms with their actual causes?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- What distortions do automated benchmarks introduce compared to real tasks?
- Do weight-space skills lose detail compared to textual skill descriptions?
- Why do generic skill descriptions evolve into execution-oriented ones?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- Why do evolved harness edits mostly memorize rather than generalize?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does training data format shape reasoning strategy more than domain?
What explains why models trained on multiple-choice data reason differently than those trained on free-form text? The research isolates format and domain effects to measure which one matters more.
format > domain at 7.5x; this adds format > instruction semantics
-
Can models pass tests while missing the actual grammar?
Do language models succeed on grammatical benchmarks by learning surface patterns rather than structural rules? This matters because correct outputs may hide reliance on shallow heuristics that fail on novel structures.
same mechanism in linguistic domain
-
Can small models reason well by just learning output format?
Does reasoning performance depend primarily on adapting how models express outputs rather than acquiring new knowledge? The Tina research tests this by applying LoRA to a 1.5B model during reasoning training.
LoRA as format adapter aligns with IT as format teacher
-
Does supervised fine-tuning actually improve reasoning quality?
While SFT boosts final-answer accuracy, does it degrade the quality and informativeness of the reasoning steps that justify those answers? This matters for high-stakes domains requiring auditable decision-making.
SFT raises accuracy because it teaches the output format, not because it improves reasoning
-
Why do chain-of-thought examples fail across different conditions?
Chain-of-thought exemplars show surprising sensitivity to order, complexity level, diversity, and annotator style. Understanding these brittleness dimensions could reveal what makes reasoning prompts robust or fragile.
complementary evidence of format-over-substance: IT achieves accuracy through format matching alone, while CoT exemplar brittleness shows reasoning performance depends on surface exemplar properties (order, style, complexity) rather than semantic content
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning
- A Survey on Post-training of Large Language Models
- Exploring Format Consistency for Instruction Tuning
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- LESS: Selecting Influential Data for Targeted Instruction Tuning
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- Instruction Induction: From Few Examples to Natural Language Task Descriptions
Original note title
instruction tuning teaches output format distribution not task understanding — simplified and delusive instructions achieve comparable performance