SYNTHESIS NOTE
Topics›Tool Computer Use›this note

Can small models match large models on function calling?

Explores whether small language models fine-tuned with the right training method can achieve comparable performance to large models on structured reasoning tasks requiring precise function calls, and what training approach makes this possible.

Synthesis note · 2026-05-03 · sourced from Tool Computer Use

The insight in this paper is methodological: function-calling for reasoning tasks is a domain where DPO outperforms SFT for small models, because the failure modes are more about preferring the right format and call sequence than about generating any plausible text. The proposed framework uses an agent that, given a problem and a callable function set, queries a large LLM by injecting function descriptions and examples and managing calls in a step-by-step reasoning chain. The byproduct is a dataset of correct AND incorrect chat completions — preference pairs ready for DPO.

Why DPO rather than SFT or PPO. SFT teaches the model to imitate good examples but provides no signal about what to avoid — and rigid output formats (precise variable names, JSON, argument values) punish near-misses harshly, so explicit negative examples matter. PPO would work but requires extensive human feedback to train a reward model, making it resource-intensive. DPO removes the reward-model step by incorporating preferences directly into the training objective, with demonstrated stability advantages over PPO.

The structural move is that a large LLM does double duty: it generates the candidate reasoning chains AND its successes/failures provide the preference labels for the small model's training. This is a teacher-distillation pattern but with both polarities — the small model learns what the large model gets right and what it gets wrong, not just to imitate the large model's right answers. The pattern fits the broader case for Can small language models handle most agent tasks?: function-calling is exactly the kind of repetitive, scoped, format-rigid work where a fine-tuned small model can replace a large general-purpose one.

The practical implication: when output format is rigid and small-model deployment is the goal, the question is not "can SFT close the gap" but "what's the cheapest source of preference signal." Self-generated preference pairs from a strong teacher are essentially free relative to human feedback.

Inquiring lines that read this note 119

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does alignment training create genuine alignment or just output compliance? Can iterative DPO replicate online reinforcement learning dynamics for research? How does decomposing tasks improve reasoning and prevent failure propagation? Can inference-time compute effectively substitute for model scale? Can prompt-based context override biases that were embedded during pretraining? Do structural constraints outperform deep architectures in recommendation systems? Does encoded knowledge in language models actually influence their outputs? Can self-generated feedback reliably guide model training without ground truth? How do training data properties determine the emergence of internal misalignment? Can language models build genuine grounding through interaction? How much do training data properties shape model reasoning? What compositional reasoning failures limit large language models despite scale? How effectively can language models perform reasoning, especially combined with symbolic methods? What training data selection strategies maximize generalization across difficulty levels? Is reasoning capability latent in base models or created by post-training? What capability trade-offs arise from domain specialization through fine-tuning? Can models improve accuracy without degrading reasoning quality? Why can't prompting alone inject genuinely new knowledge into models? Why do stronger reasoning capabilities create tradeoffs with instruction following? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? How do capability benchmark scores systematically misrepresent true model abilities? What training dynamics and scale trigger emergence of reasoning capabilities? How do evaluation practices shape which failures stay visible? How should inference compute be allocated based on problem difficulty? Can intelligent routing over smaller models outperform scaling a single large model? What causes retrieval-augmented generation systems to fail despite access to external knowledge? What reasoning architectures enable models to solve complex problems efficiently? How much does training format versus domain influence reasoning? What role does sparsity play in model behavior and scaling decisions? What makes distillation transfer some model capabilities while suppressing others? How do surface patterns enable correct outputs but reduce robustness? Why does adding new knowledge through fine-tuning degrade existing capabilities? Does RL create genuinely new reasoning capabilities or refine existing ones? What enables genuine semantic understanding in language models? Why do embedding systems fail to capture task-relevant relationships? Is language model reasoning authentic and what causes models to reason? How does the generation-verification gap limit what we can measure about AI reasoning? What causes reasoning models to fail or wander off track? How do multi-agent LLM systems fail distinctly compared to single agents? What fundamental constraints limit how effectively agents can improve themselves? Should GUI agents use structured representations over raw visual input? How does harness optimization generalize across different model architectures and domains? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Can compression size predict model complexity better than parameter count alone? How do standardized protocols improve multi-agent coordination and reliability? When do multi-agent systems outperform single frontier models? How should agents manage memory granularity to improve long-term performance? How should designers communicate what AI systems truly are and can do? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 152 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

DPO-trained small models can match large models on function-calling reasoning chains — preference data from a teacher beats SFT for the rigid output format