Can genuine reasoning activation coexist with contaminated benchmarks?
RLVR shows both real behavioral changes and inflated metrics. Can these contradictory findings actually describe the same phenomenon from different angles, and what does that mean for evaluating reasoning improvements?
Two RLVR findings appeared contradictory:
Spurious rewards work: Why do random rewards improve reasoning for some models but not others?, suggesting the reward signal itself matters less than the RL training process, which activates latent pretraining capabilities. This was treated as evidence that RLVR functions as a pretraining catalyst rather than a reasoning teacher.
Benchmark contamination: Since Does RLVR success on math benchmarks reflect genuine reasoning improvement?, the metric improvement may be data memorization rather than genuine reasoning activation.
The resolution: These findings operate at different measurement levels and can coexist:
Behavioral activation (genuine): RL training with any reward signal activates code reasoning formats and structured thinking patterns that exist in pretraining data but are dormant. This is visible in output format changes, thinking token usage, and exploration behavior changes — measurements not contaminated by benchmark overlap.
Benchmark improvement (inflated): The metric improvement on contaminated benchmarks is partially or fully attributable to memorization. Clean benchmarks show reduced or eliminated gains for spurious rewards, while correct rewards still improve.
The practical implication: RLVR research must separate behavioral measurements (how the model's reasoning process changes) from performance measurements (how benchmark scores change). Both are informative; conflating them produces confusion about what RLVR actually does. The one-shot activation finding (single example triggers 36%→73.6% improvement) may itself need re-evaluation on clean benchmarks.
Inquiring lines that read this note 77
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we distinguish genuine model deception from honest errors?- What makes accountability and validity-orientation non-behavioral properties?
- How can we detect dishonesty in model outputs separate from capability failures?
- What other hidden biases might aggregate metrics fail to distinguish from reasoning?
- How should we redesign benchmarks to catch conservative bias in reasoning tasks?
- How much RLVR improvement comes from benchmark data memorization?
- Why do benchmark designers treat content effects as confounds?
- How much do metric choices inflate claims about model capabilities?
- How do weight perturbations reveal what performance benchmarks cannot measure?
- Why do benchmark scores not capture the true nature of AI systems?
- What deployment context determines which benchmark mode actually matters?
- What evaluation methods actually measure reasoning versus execution capability?
- How much of MATH-500 improvement comes from data contamination versus real reasoning gains?
- What training regimes confound surface mechanisms with their actual causes?
- Why do AI benchmarks show rapid saturation from near-zero to near-perfect?
- Why do static benchmarks miss frontier capabilities that open-world tasks reveal?
- What real-world tasks most clearly expose gaps between benchmark performance and actual capability?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- Why do benchmarks become saturated so quickly after initial launch?
- How much improvement comes from caching versus actual capability gain?
- Can an average-case validator score hide poor performance on critical tasks?
- How does measurement error in capability benchmarks systematically underestimate or overestimate true ability?
- How should benchmarks balance verifiability against outcome resolution?
- Why do hidden test partitions matter more than open evaluation sets?
- Why does benchmark saturation give a false sense of capability coverage?
- When does measured progress on an evaluator conceal actual performance decline?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- What distortions do automated benchmarks introduce compared to real tasks?
- Can clean benchmarks reveal true RLVR reasoning gains?
- Should benchmarks measure trace length or whether constraints were actually satisfied?
- How does task contamination differ from test set data leakage?
- How does tool access change what we measure in reasoning tests?
- Why do current RLVR methods fail to expand reasoning capability beyond base model boundaries?
- Are RLVR models worse than non-reasoning models for subjective annotation?
- Does RLVR expand model capability or reorganize existing capability?
- Why does medium difficulty outperform both easy and hard RLVR training samples?
- Does RLVR teach new reasoning or activate existing pretraining capabilities?
- Can combining SRL with RLVR outperform either method used alone?
- How does RLSVR differ from using model probability or self-judgment?
- Can reasoning benchmarks separate logic from believability?
- Can activation patching reveal which reasoning steps actually matter?
- Why do benchmark scores rise while reasoning quality declines?
- Can benchmark improvements hide degradation of deliberative reasoning?
- How do surface correlations between narratives and answers mislead benchmark validity?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- How do live human evaluations differ from ground-truth benchmarks?
- Why do open-world evaluations reveal capabilities that static benchmarks hide?
- How much reasoning catalyst data is actually needed for improvement?
- How does RPT compare to learning when versus how to deploy reasoning?
- What pretraining formats encode latent reasoning strategies that RLVR can surface?
- Why do invalid reasoning steps produce nearly the same performance gains?
- What distinguishes genuine capability gains from coherent but invalid reasoning traces?
- Does RLVR reward structure create pressure toward traces that look right?
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- How do satisfaction scores differ from genuine cognitive improvement?
- Why does accumulated portfolio output not match accumulated worker capability?
- What makes single-axis benchmarks systematically misrepresent deployment readiness?
- How do benchmark environments misrepresent deployment readiness?
- What realism standards should compliance benchmarks meet to avoid evaluation gaming?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- Spurious Rewards: Rethinking Training Signals in RLVR
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
- Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Original note title
RLVR behavioral activation and benchmark improvement are separable — genuine pretraining activation can coexist with contamination-inflated metrics