SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Do fine-tuned language models actually learn optimization procedures?

Can RL fine-tuning teach LLMs to solve constraint-optimization problems through genuine reasoning, or does it merely sharpen pattern-matching? Testing on out-of-distribution variants reveals the mechanism.

Synthesis note · 2026-05-18 · sourced from Reasoning Architectures

The constraint-optimization study uses a clean diagnostic to separate procedure from pattern: an N-case test set (in-distribution power-grid topologies) and an N-1 test set (the same problems with one element removed, putting them out of distribution while keeping the structure recognizable). A model running the actual procedure should perform comparably on both. A model running pattern-match should perform worse on N-1.

Even under GRPO and constraint-satisfaction-reward training, models degrade markedly on N-1. The conclusion is that RL on outcome-based rewards does not install the missing procedure — it sharpens the template-matching strategy along the in-distribution axis. The model gets better at recognizing patterns it has seen and worse, relatively, at adapting to perturbed structure.

This is methodologically important because it provides a probe that other reasoning evaluations lack. Most benchmarks cannot distinguish "the model solved this" from "the model recognized this." The N / N-1 comparison forces the distinction by holding the problem class fixed while perturbing the instance. The drop is the memorization signature.

For practitioners, the diagnostic generalizes. Wherever a deployment cares whether a model is computing or recalling — clinical reasoning, legal-statute reasoning, scientific problem-solving — building an "N-1" counterpart of the canonical test set is a cheap way to surface memorization. The structure-shift probe is more informative than headline accuracy on the canonical set.

Inquiring lines that read this note 103

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does alignment training create genuine alignment or just output compliance? How do surface patterns enable correct outputs but reduce robustness? Do reasoning benchmarks predict model performance in long-horizon workflows? Why do LLM recommenders underperform collaborative filtering despite their capabilities? How do capability benchmark scores systematically misrepresent true model abilities? How do neural networks achieve compositional generalization at scale? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do training data properties determine the emergence of internal misalignment? What capability trade-offs arise from domain specialization through fine-tuning? Why can't prompting alone inject genuinely new knowledge into models? What compositional reasoning failures limit large language models despite scale? Does RL create genuinely new reasoning capabilities or refine existing ones? How can evolutionary algorithms maintain diversity during solution search? Can models improve accuracy without degrading reasoning quality? Why don't LLMs reliably translate capability into accurate outputs? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What makes distillation transfer some model capabilities while suppressing others? How does improved reasoning affect models' ability to acknowledge uncertainty? Can prompt-based context override biases that were embedded during pretraining? Do language models learn genuine understanding or just surface patterns? How much do training data properties shape model reasoning? Is language model reasoning authentic and what causes models to reason? Why does adding new knowledge through fine-tuning degrade existing capabilities? How effectively can language models perform reasoning, especially combined with symbolic methods? What enables genuine semantic understanding in language models? How do standardized protocols improve multi-agent coordination and reliability? What training data selection strategies maximize generalization across difficulty levels? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Can memory architectures handle ultra-long context better than attention? Does preference optimization systematically degrade conversational grounding in language models? How do pretraining biases affect reward signal effectiveness in RLVR? What training dynamics and scale trigger emergence of reasoning capabilities? Is reasoning capability latent in base models or created by post-training? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Does abstract user knowledge outperform concrete interaction history in personalization? What is the relationship between thinking tokens and reasoning accuracy?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 109 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

N-1 out-of-distribution tests reveal that RL fine-tuned LLMs still rely on memorization for optimization problems