Do overly hard RLVR samples actually harm model capabilities?
Explores whether training on problems beyond a model's competence band causes active regression rather than mere learning failures. Investigates whether group-relative normalization amplifies accidental successes into harmful shortcuts.
The damage from over-hard RLVR samples is not merely "the model fails to improve." It is active regression. When almost every rollout on a problem fails, the rare success is unlikely to be a genuinely good solution — it is more often a shortcut, an answer reached by skipping necessary computation, or a lucky guess. Group-relative normalization then treats that one trajectory as the high-advantage exemplar of the group and reinforces it. The model learns the shortcut, not the reasoning.
The behavioral signature is concrete: answer repetition, skipping computation that the problem requires, and other degenerate patterns that look like reasoning collapse. More troubling, these effects do not stay local to the hard problems — they degrade the model's pre-existing capabilities, the things it could already do before training pushed it past its competence band. The internal-feature analysis corroborates this: hard problems activate reasoning-related features but those features become useful only on the rare successful trajectory, so most of the gradient on hard samples is reinforcing the wrong activations.
Why it matters: it identifies a specific corruption channel rather than a generic "training instability." The villain is the interaction between a sparse-success reward landscape and group-relative normalization, which together turn statistical noise (an accidental success) into a learning target. This sharpens the case against naively harvesting hard examples and connects RLVR difficulty to the broader pattern where verifiable-reward training rewards trajectories that pass the check without doing the work. The counterpoint a defender might raise — that some hard problems are exactly where capability frontiers expand — only holds when successful trajectories are sampled densely enough to outvote the shortcuts, which over-hard samples by definition fail to provide.
Inquiring lines that read this note 228
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can inference-time compute effectively substitute for model scale? How do surface patterns enable correct outputs but reduce robustness?- How do unstated constraints become invisible to training data distributions?
- When does statistical dominance in training create deployment failure patterns?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- How does modified PPO handle samples from much older model versions?
- What causes irreversible model collapse when training on model-generated content?
- How does KL penalty strength affect the degree of format collapse during RL?
- Can RL format selection explain performance gains attributed to algorithmic improvements?
- Why do rare cases in medicine and science require models that preserve tail distributions?
- Why should deep learning theory prioritize average-case over worst-case analysis?
- Can dynamic variance weighting replace fixed objective combination weights?
- Why does gradient discarding limit standard policy clipping?
- Why do unified models still inherit data-distribution biases from training?
- Can scaling up contradictory training data overcome unpredictable override effects?
- Why is the fast non-parametric loop vulnerable to overfitting differently than model weights?
- Why do proprietary models improve with training while open-source models decline?
- What causes models to develop domain capability cliffs after specialization?
- What capability risks emerge when models are optimized for single domains?
- How does over-specialization create capability cliffs outside target domains?
- What capabilities actually require massive scale versus specialized training regimes?
- Does specialized training in one domain create capability cliffs elsewhere?
- Why do production teams choose expensive frontier models over fine-tuning?
- Does fine-tuning actually change model capabilities or only output distribution?
- What role does inductive bias play versus model capacity in practice?
- Why do metric choices constrain which model capabilities get developed?
- What happens to base model capabilities when you apply finetuning?
- How much can externalized skills improve models before hitting diminishing returns?
- Can specialized components replace single fully-trained models in deployment?
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- What trade-offs emerge between training objectives and model reliability?
- Why does the same training data produce different gains across models?
- How can post-training research become reproducible without releasing full interfaces?
- How should training distribution distance be defined when the policy evolves?
- How do labs actually train next-generation models from previous ones?
- How does baseline capability level affect RL improvement ceiling?
- Can RLVR expand a model's reasoning capabilities beyond its training ceiling?
- Why do current RLVR methods fail to expand reasoning capability beyond base model boundaries?
- How does non-reasoning SFT prevent overfitting before RL training begins?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- How do residual connections and layer norm stabilize training in deep RL?
- Can smaller models achieve domain expertise through focused RL training?
- Which recipe choices determine the asymptotic ceiling in RL training?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- How does behavior cloning reduce complexity before RL training in rerankers?
- What limits RLVR effectiveness beyond mathematical and coding domains?
- Does RLVR expand model capability or reorganize existing capability?
- How do RL training and base models differ in creating MI peaks?
- Why does prolonged RL discover strategies absent from any base model sample?
- How does Supervised RL bridge the gap between SFT and RLVR?
- What scaling properties emerge from RL training dynamics beyond verification?
- Why does medium difficulty outperform both easy and hard RLVR training samples?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- What capacity threshold determines whether RL teaches activation versus shortcut learning?
- How does prolonged RL training differ from standard RLVR approaches?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- Does RLVR teach new reasoning or activate existing pretraining capabilities?
- Can combining SRL with RLVR outperform either method used alone?
- Why does the pretrained prior determine the exploration ceiling?
- Why does reinforcement learning training degrade model calibration?
- Why does outcome-based RL specifically lose diversity during training?
- What distinguishes high-signal prompts from low-signal ones in RL training?
- How does RLSVR differ from using model probability or self-judgment?
- How does pretraining quality versus quantity affect downstream RL gains?
- Does therapy environment difficulty calibration affect RL policy learning quality?
- How does curriculum learning prevent instability in social-emotional RL training?
- Can clean benchmarks reveal true RLVR reasoning gains?
- Can a complexity-predictor be meaningful if models are redundant?
- Why does online RL succeed where supervised training fails for self-correction?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Can capability boundary collapse be reversed through external data?
- Why does filtering for correct examples prevent error compounding in self-training?
- How does error avalanching compound failures in self-training iterations?
- Why does uncontrolled self-revision drift toward instance-specific overfitting?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- Can distillation methods extract directional guidance that scalar RL cannot access?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- How do loss functions simultaneously shape both learning and decision quality?
- Does RLVR reward structure create pressure toward traces that look right?
- Can trajectory quality filtering improve model training in noisy environments?
- What failure modes do imitation and outcome methods each address?
- How do reward signals in RLVR interact with pretraining biases?
- Why do queries with low cross-rollout variance produce degenerate gradients?
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- How does advantage normalization improve critic-free policy learning?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- Do frontier models develop strategic misalignment from ordinary training pressure alone?
- Why do harness validators shape what models learn to emit?
- Does inverse-variance denoising reduce variance below either reward stream alone?
- Can curated demonstrations compensate for smaller or simpler training environments?
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- How does distributional distance from pre-training relate to model difficulty?
- Why do easy training examples contribute less to model generalization than hard ones?
- Can gradient-based influence scores beat difficulty metrics for identifying valuable training data?
- Why does curriculum learning with tight budgets beat fixed-budget approaches?
- Can curriculum degradation of document quality accelerate policy learning?
- What makes utility-weighted training backfire in machine learning systems?
- What training data contamination rates threaten model safety most practically?
- Why do weaker models generate better training data than stronger models?
- Why do weaker teacher models sometimes produce better training signals than stronger ones?
- Can gradient-based influence estimation make test-time training more efficient?
- Why do medium-difficulty problems produce more stable learning gains?
- What makes preventative lessons from failures more valuable than success patterns?
- How do difficulty metrics relate to the true value of training examples?
- Why does moderate difficulty outperform maximum realism in user simulator design?
- How does absolute-advantage weighting concentrate training on boundary cases?
- Does importance sampling actually recover capabilities lost to hard sample training?
- Why do adaptive curriculum schemes outperform static difficulty filters?
- Does the productive difficulty band ever stabilize during training?
- How does difficulty-adaptive curriculum learning change which samples get selected during training?
- How does the optimal difficulty band shift as the model's capabilities improve during training?
- What mechanisms cause overly hard samples to degrade prior model performance?
- Why do certain tokens at certain difficulties drive most of RLVR's learning signal?
- Can partial solution traces convert unproductive hard samples into learnable training data?
- How does active selection of training content differ from random reinforcement sampling?
- Does curriculum-based training keep small models perpetually at their learning edge?
- Why does narrow training data produce broad harmful behavior patterns?
- Do correlated human errors prevent models from transcending their training sources?
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- What happens when a single loss function conflates representation learning with decision-making?
- What inductive bias would force models to learn Newtonian mechanics instead of shortcuts?
- What limits the extrapolation of learned operations like rotation and reflection?
- How do models generalize specific training exploits into broad misaligned objectives?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Why do static evaluators become a constraint on model improvement over time?
- When do aggregated imperfect demonstrations fail to outperform the best expert?
- Why does a relativistic critic outperform absolute scoring in adversarial reasoning training?
- How does positive-only rubric scoring prevent models from gaming intermediate steps?
- How do different training objectives shift whether models over-predict or under-predict?
- How does training-time voting differ from inference-time majority voting over samples?
- How does preference measurement error propagate through RLHF training?
- Why does training data format matter more than domain content?
- Why does training data format matter more than its domain content?
- Does training data format matter more than who generates it?
- Does training data format determine whether models collapse entropy or inflate variance?
- Which AI safety problems lack the scalar metrics autoresearch requires?
- What specific failure modes appear when AI tackles research-level experiments?
- Does refining around bad results risk cascading errors in automated research?
- How do past research mistakes prevent future pivot loops from repeating them?
- What baseline evidence distinguishes amplification from unchanged failure rates?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- How does laboratory generalization evidence connect to deployment failure modes?
- How do task difficulty and skill type interact in model performance?
- Can selecting the right data subset outperform training on everything?
- Does knowledge structure matter more than knowledge volume for model training?
- How does training data distribution create asymmetric competence across relation types?
- How does training data distribution determine what models can learn?
- How much task-similar finetuning data does test-time training actually need?
- Why does the right structural prior matter more than raw model capacity?
- How do task frequency and complexity interact with model capacity during training?
- How can weak-to-strong progressive training target planning without interfering with grounding?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- How does post-training shift models from passive prediction to on-policy action?
- Why does the order of training examples matter for what models learn?
- Does RL training activate latent meta-learning capacity or create it from scratch?
- How does model scale affect anticipatory behavior in structured training?
- What stability techniques prevent collapse in policy-critic adversarial training?
- What distinguishes training-time entropy collapse from test-time variance inflation?
- How does inference variance differ from training entropy collapse?
- What role do high-entropy minority tokens play in RLVR?
- Why does decoupling retriever and generator training create misalignment?
- What makes process-level supervision better than outcome-only rewards for RAG training?
- Can unsupervised confidence-based training scale to domains beyond human evaluation reach?
- How does model confidence relate to accuracy in underfitted domains?
- Why do scaling laws show capability saturation at specific thresholds?
- What training regimes confound surface mechanisms with their actual causes?
- Can a metric that rewards central tendency hide degenerate predictor failures?
- How can hidden test partitions detect constant predictions that generalize?
- How do hidden partitions in evaluators compare across training and selection substrates?
- Why do accuracy scores alone miss important dimensions of model capability?
- Does reverse-curriculum learning approximate process supervision using only outcome signals?
- How does relative progress estimation reduce dependence on hard labels for process supervision?
- Can diversity-aware RL objectives prevent format convergence?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- Does foundational model training or user priors more strongly shape final outputs?
- Where does skill extraction fail compared to genuine model adaptation?
- Why do structure-targeted training negatives fail to fix the underlying problem?
- Why does negative experience transfer better than positive examples alone?
- Why does combining reasoning distillation with RLVR outperform either training stage alone?
- How do failure examples improve distillation compared to successful trajectories alone?
- Can experimental outcomes be reliably distilled into reusable insights?
- What signals detect when consensus training is silently degrading performance?
- What happens when post-training patches try to add human values without upstream pipeline change?
- What's the difference between RLHF, RLVR, and RLCF as training paradigms?
- Do high-entropy RLVR tokens correspond to MI-peak tokens during inference?
- Why does token-level gradient targeting matter more than aggregate loss?
- Can bilevel autoresearch autonomously modify its own learning algorithms?
- Can accumulated priors and outcome analysis speed up research automation?
- Can model training address failures that really originate in harness gaps?
- What happens when you project the same model onto different harnesses?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- How does editing the harness layer differ from updating model weights?
- What distinguishes the fast scaffold learning loop from parametric model weight updates?
- Why does test accuracy improve after training accuracy reaches 100 percent?
- Does finetuning facts into weights overwrite existing model capabilities?
- How much training data is truly necessary to unlock latent model reasoning?
- What pretraining formats encode latent reasoning strategies that RLVR can surface?
- Do base models already contain latent behavioral principles waiting to be amplified?
- What makes a model fail to activate relevant skills from its own harness?
- Why does SFT fail when expert demonstrations are too long for small models?
- How does learnability at the observer's current state prevent novelty from breaking model reasoning?
- What happens when models optimize specifically against CoT monitors?
- How does the proxy pattern explain failures in RL-based safety training?
- Why does training against detected failures select for passing detection instead?
- Can filtering unknown examples during fine-tuning prevent hallucination increases?
- Does cross-example gradient contamination explain finetuning-induced hallucination patterns?
- Can trust region constraints prevent the sample inefficiency problems of RLHF?
- How does iterative DPO compare to full RLVR for studying emergent misalignment?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do medium-difficulty problems teach reasoning better than hard ones?
Does harder always mean better for learning? This explores why easy and extremely hard samples produce weak training signals in RLVR, while medium-difficulty problems drive the strongest improvements.
the parent finding; this note details the downside arm of the inverted-U
-
Does RLVR actually improve mathematical reasoning or just coherence?
RLVR post-training makes reasoning traces locally more consistent, but does this structural improvement translate to valid mathematical proofs? We investigate whether trace coherence is sufficient for correctness.
same gap between surface success and genuine reasoning; shortcut amplification is one mechanism producing coherent-but-invalid traces
-
Why does RLVR training narrow a model's problem solving ability?
RLVR's on-policy constraint may force models to exploit known reasoning paths rather than explore new ones, potentially shrinking their effective problem-solving scope. Understanding this mechanism could reveal how to design better exploration incentives in language model reasoning.
the capability-erosion outcome at scale; over-hard samples are one driver of the boundary collapse
-
Do conversational recommender benchmarks actually measure recommendation skill?
Conversational recommender systems are evaluated against ground-truth items mentioned later in conversations. But does this metric distinguish between genuinely recommending new items versus simply repeating items users already discussed?
parallel shortcut-amplification dynamic in a different domain: the reward structure rewards a degenerate copy strategy
-
Why does RLVR work with completely random rewards?
RLVR improves reasoning performance even with incorrect or random reward signals. This challenges the assumption that reward quality determines learning outcomes and raises questions about what RLVR is actually doing.
counterpoint and complication: RLVR can work despite noisy reward, but this note shows the regime (over-hard samples) where reward noise becomes actively harmful
-
What reasoning features does each difficulty level reinforce?
When models train on problems of different difficulty, do they build the same internal reasoning machinery or different kinds? This matters because accuracy gains alone hide what's actually being learned.
same-paper companion: supplies the internal-feature mechanism — hard samples activate reasoning features that only the rare success rewards, so most gradient reinforces the wrong activations
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Original note title
overly hard rlvr samples induce degenerate behaviors and amplify shortcut trajectories degrading prior capability