SYNTHESIS NOTE
Topics›Flaws›this note

Can genuine reasoning activation coexist with contaminated benchmarks?

RLVR shows both real behavioral changes and inflated metrics. Can these contradictory findings actually describe the same phenomenon from different angles, and what does that mean for evaluating reasoning improvements?

Synthesis note · 2026-02-23 · sourced from Flaws

Two RLVR findings appeared contradictory:

Spurious rewards work: Why do random rewards improve reasoning for some models but not others?, suggesting the reward signal itself matters less than the RL training process, which activates latent pretraining capabilities. This was treated as evidence that RLVR functions as a pretraining catalyst rather than a reasoning teacher.

Benchmark contamination: Since Does RLVR success on math benchmarks reflect genuine reasoning improvement?, the metric improvement may be data memorization rather than genuine reasoning activation.

The resolution: These findings operate at different measurement levels and can coexist:

  1. Behavioral activation (genuine): RL training with any reward signal activates code reasoning formats and structured thinking patterns that exist in pretraining data but are dormant. This is visible in output format changes, thinking token usage, and exploration behavior changes — measurements not contaminated by benchmark overlap.

  2. Benchmark improvement (inflated): The metric improvement on contaminated benchmarks is partially or fully attributable to memorization. Clean benchmarks show reduced or eliminated gains for spurious rewards, while correct rewards still improve.

The practical implication: RLVR research must separate behavioral measurements (how the model's reasoning process changes) from performance measurements (how benchmark scores change). Both are informative; conflating them produces confusion about what RLVR actually does. The one-shot activation finding (single example triggers 36%→73.6% improvement) may itself need re-evaluation on clean benchmarks.

Inquiring lines that read this note 77

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we distinguish genuine model deception from honest errors? How do capability benchmark scores systematically misrepresent true model abilities? What design and behavioral factors drive false consciousness attribution to AI? Do reasoning benchmarks predict model performance in long-horizon workflows? Why is hallucination an inevitable limitation of current language models? How can we prevent synthetic data from contaminating statistical inference and corpora? Does RL create genuinely new reasoning capabilities or refine existing ones? Can models improve accuracy without degrading reasoning quality? How does the generation-verification gap limit what we can measure about AI reasoning? Is reasoning capability latent in base models or created by post-training? How well do AI systems understand human social norms? Do reasoning traces faithfully reflect actual model reasoning? What makes distillation transfer some model capabilities while suppressing others? Should agents decouple planning from perception grounding for better performance? How do pretraining biases affect reward signal effectiveness in RLVR? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can local safety checks guarantee system-level behavioral safety? Why does polished presentation create unearned authority in AI outputs? What trajectory-level metrics beyond task success best evaluate agent performance? What training data selection strategies maximize generalization across difficulty levels? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? How should agent systems validate and persist generated code artifacts? Why do standard benchmarks fail to predict agent deployment success? How can infrastructure records verify actual agent behavior? What capability trade-offs arise from domain specialization through fine-tuning? How do false presuppositions and sycophancy drive persistent false beliefs in models? How do surface patterns enable correct outputs but reduce robustness? What should agent evaluation prioritize to reveal reliable behavior?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

RLVR behavioral activation and benchmark improvement are separable — genuine pretraining activation can coexist with contamination-inflated metrics