SYNTHESIS NOTE
Topics›Test Time Compute›this note

Can non-reasoning models catch up with more compute?

Explores whether inference-time compute budget can close the performance gap between standard models and those trained for reasoning, and what training mechanisms might enable this.

Synthesis note · 2026-02-20 · sourced from Test Time Compute

In verifier-free inference-time compute experiments (Think Deep, Think Fast), non-reasoning models fall substantially behind reasoning models even when given an extremely high inference budget. The gap doesn't close with more compute — it just stays there.

This sets a hard limit on Can inference compute replace scaling up model size?. The substitution works within a training regime, but not across training regimes. A standard instruction-tuned model with more inference compute cannot replicate what a model trained specifically for extended reasoning can do, even given equivalent token budgets.

Why? Reasoning models have internalized the reasoning process through training — they know how to use additional tokens productively. Non-reasoning models don't have this structure, so additional tokens degrade into noise or verbosity rather than improved reasoning. The training regime instills the reasoning protocol that makes inference compute usable.

Qualification from targeted activation (Base Models paper): The gap is substantially closeable through targeted steering of base model activations without weight updates. A hybrid model using base model weights + thinking model deployment decisions recovers 91% of the performance gap while steering only 12% of tokens. This doesn't invalidate the finding — non-reasoning models without steering still fall behind — but it significantly changes what "non-reasoning model" means in practice. If capability already exists latently and steering can surface it, the gap is about deployment mechanisms, not raw capability. See Does RL teach reasoning or just when to use it?.

The imitation learning ceiling (Tutorial on LLM Reasoning): SFT/imitation learning creates an intelligence upper bound: the model is bounded by the quality of demonstrations it learns from, unable to surpass the skill level present in training data. RL + world models is the path beyond this ceiling, because RL allows discovery of strategies that exceed any individual demonstration. This provides the mechanism for why reasoning-specific training matters: it is not merely "more training" but training that enables exceeding the imitation ceiling.

This is a strong argument for the necessity of reasoning-specific post-training, not just inference-time tricks. Compute can amplify capability but cannot manufacture it. The dependency on training regime appears to be capability-specific: Can language models learn grammar from child-scale data? — syntactic competence scales down readily, achievable with human-scale data and the right composition. Reasoning capability requires the opposite: specialized training that instills the reasoning protocol itself. The lesson is not "you need a bigger model" but "you need the right training for the capability you want."

Inquiring lines that read this note 198

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can models improve accuracy without degrading reasoning quality? What training dynamics and scale trigger emergence of reasoning capabilities? How does decomposing tasks improve reasoning and prevent failure propagation? Do structural constraints outperform deep architectures in recommendation systems? Can reasoning scale in latent space without tokens? How should inference compute be allocated based on problem difficulty? What reasoning architectures enable models to solve complex problems efficiently? Can intelligent routing over smaller models outperform scaling a single large model? Can inference-time compute effectively substitute for model scale? Can parallel reasoning outperform sequential reasoning under fixed token budgets? What is the relationship between thinking tokens and reasoning accuracy? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Can diffusion models match autoregressive performance on language generation tasks? Why does adding new knowledge through fine-tuning degrade existing capabilities? How should test-time compute scaling work in agentic systems? Does model confidence reliably signal actual accuracy in practice? How does reasoning length affect model performance across different tasks? What compositional reasoning failures limit large language models despite scale? How do prompting refinements mask underlying biases and model frequency patterns? What causes reasoning models to fail or wander off track? Why do stronger reasoning capabilities create tradeoffs with instruction following? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What causes retrieval-augmented generation systems to fail despite access to external knowledge? When do multi-agent systems outperform single frontier models? What role does sparsity play in model behavior and scaling decisions? What attack surfaces do reasoning traces and chains introduce? How effectively can language models perform reasoning, especially combined with symbolic methods? Can memory architectures handle ultra-long context better than attention? Why do token-level mechanisms matter for learning to reason? Is reasoning capability latent in base models or created by post-training? How does policy entropy collapse constrain scaling of reasoning-focused RL? How should systems decide whether to retrieve or reason alone? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How do neural networks achieve compositional generalization at scale? Do reasoning traces faithfully reflect actual model reasoning? How do pretraining biases affect reward signal effectiveness in RLVR? How does the generation-verification gap limit what we can measure about AI reasoning? How do surface patterns enable correct outputs but reduce robustness? Do reasoning benchmarks predict model performance in long-horizon workflows? Can prompt-based context override biases that were embedded during pretraining? What makes step-level supervision effective for complex reasoning traces? Does RL create genuinely new reasoning capabilities or refine existing ones? How does synthetic data quality and diversity affect downstream model capabilities? How much do training data properties shape model reasoning? Do language models develop actual world models or merely task heuristics? What makes distillation transfer some model capabilities while suppressing others? Can we reliably detect when models game evaluations? How can evolutionary algorithms maintain diversity during solution search? How should designers communicate what AI systems truly are and can do? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
22 direct connections · 234 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

non-reasoning models cannot match reasoning models even with unlimited inference budget