SYNTHESIS NOTE
Topics›Linguistics, NLP, NLU›this note

Are models actually reasoning about constraints or just defaulting conservatively?

Do language models genuinely apply constraints when solving problems, or do they simply prefer harder options by default? Minimal pair testing reveals whether apparent reasoning success masks hidden biases.

Synthesis note · 2026-05-01 · sourced from Linguistics, NLP, NLU

The Heuristic Override Benchmark uses minimal pairs — same surface heuristic, with versus without the implicit constraint — to test whether apparent reasoning successes reflect actual reasoning. The result is striking. Twelve of fourteen models perform worse on the no-constraint variant than on the constraint-active variant, with drops up to 38.5 percentage points. Only two models (GPT-OSS-120B at +13.8 and GPT-OSS-20B at +11.0) improve when the constraint is removed.

This exposes a hidden mechanism behind apparent accuracy. When the constraint is present, the correct answer is the harder one (drive to the car wash that is 50m away). When the constraint is removed, the correct answer is the easier one (walk to the store that is 50m away). Models that default to recommending the harder option score correctly on constraint-active cases without doing any constraint reasoning. They are not solving the problem. They are reflexively choosing the more conservative option, which happens to coincide with the constraint-required answer.

The minimal-pair asymmetry is the only test that catches this. Single-instance accuracy looks fine — the model recommended driving, the right answer was driving. But the same model recommends driving even when walking would be correct, because the recommendation is not based on the constraint. The two-of-fourteen models that improve on minimal pairs are the only ones whose constraint-active accuracy reflects genuine reasoning about the constraint. The rest are riding a conservative-bias accident that aggregate metrics cannot distinguish from reasoning.

Inquiring lines that read this note 125

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do stronger reasoning capabilities create tradeoffs with instruction following? What causes reasoning models to fail or wander off track? How do capability benchmark scores systematically misrepresent true model abilities? Can prompt-based context override biases that were embedded during pretraining? What reasoning architectures enable models to solve complex problems efficiently? Can inference-time compute effectively substitute for model scale? How effectively can language models perform reasoning, especially combined with symbolic methods? How do training data properties determine the emergence of internal misalignment? How do surface patterns enable correct outputs but reduce robustness? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How do prompting refinements mask underlying biases and model frequency patterns? Can models improve accuracy without degrading reasoning quality? How should designers communicate what AI systems truly are and can do? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? How does self-revision in reasoning models affect accuracy and confidence? How does improved reasoning affect models' ability to acknowledge uncertainty? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Do language models learn genuine understanding or just surface patterns? Do reasoning traces faithfully reflect actual model reasoning? Do reasoning benchmarks predict model performance in long-horizon workflows? Does encoded knowledge in language models actually influence their outputs? Is language model reasoning authentic and what causes models to reason? How do false presuppositions and sycophancy drive persistent false beliefs in models? What compositional reasoning failures limit large language models despite scale? Why don't LLMs reliably translate capability into accurate outputs? What emerges when safety-aligned models attempt to role-play deceptive personas? How does persona conditioning amplify demographic stereotyping and bias in models? Can inoculation prompting prevent emergent misalignment after reward hacking? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What is the relationship between thinking tokens and reasoning accuracy? Do language models reason through causal mechanisms or semantic associations? What capability trade-offs arise from domain specialization through fine-tuning? Why do language models resist personality conditioning through prompts? How does reasoning length affect model performance across different tasks? How does the generation-verification gap limit what we can measure about AI reasoning? What training data selection strategies maximize generalization across difficulty levels? How does evaluation scope and dimensionality affect what we measure? Does alignment training create genuine alignment or just output compliance? Should agents decouple planning from perception grounding for better performance? How do pretraining biases affect reward signal effectiveness in RLVR? How do prompt design choices influence model reasoning and performance? Why do token-level mechanisms matter for learning to reason? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Why do locally safe actions create system-level safety gaps? Can we reliably detect when models game evaluations? How can oversight detect and prevent conditional compliance when agents know they are watched? Can brute-force automated research substitute for iterative depth and human research intuition?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Conservative bias hides behind apparent reasoning success — most models perform worse when the constraint is removed than when it is present