SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Does negative reinforcement alone outperform full reinforcement learning?

Can training with only penalty signals for wrong answers match or exceed full RL approaches? This challenges the conventional assumption that reward design requires both positive and negative signals.

Synthesis note · 2026-02-22 · sourced from Reinforcement Learning

Decomposing RL's learning signal into positive sample reinforcement (PSR) and negative sample reinforcement (NSR) reveals a surprising asymmetry. Training with only negative samples — penalizing incorrect responses without ever reinforcing correct ones — consistently improves performance over the base model across the entire Pass@k spectrum (k up to 256), often matching or surpassing full PPO and GRPO.

The mechanism is straightforward through gradient analysis: NSR works by suppressing incorrect generations and redistributing probability mass toward other plausible candidates, guided by the model's prior beliefs. It refines existing knowledge rather than introducing entirely new behaviors. This is because penalizing a wrong answer doesn't point toward any specific correct answer — it lets the model's own prior determine where the freed probability mass flows.

Positive-only reinforcement creates the opposite problem. It improves Pass@1 (the model gets better at its top-ranked answer) but degrades performance at higher k because it concentrates probability mass on rewarded trajectories, reducing diversity. Since Does policy entropy collapse limit reasoning performance in RL?, positive reinforcement actively contributes to the problem while negative reinforcement sidesteps it.

This reframes how we think about RL for reasoning. The conventional framing is that RL rewards correct behavior. But the evidence suggests that penalizing incorrect behavior may contribute more to performance than reinforcing correct behavior — especially when diversity matters. The model already contains good solutions in its prior; it just needs help avoiding the bad ones.

The practical implication is that reward design for reasoning RL may be over-engineered. If suppression alone gets you most of the way, the elaborate reward shaping and process supervision architectures may be solving a problem that's already largely solved by the base model's prior distribution.

Inquiring lines that read this note 83

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Do structural constraints outperform deep architectures in recommendation systems? How do pretraining biases affect reward signal effectiveness in RLVR? Do language models lack essential therapeutic presence and engagement? Does RL create genuinely new reasoning capabilities or refine existing ones? How do spurious versus genuine rewards shape model reasoning and behavior? Can diffusion models match autoregressive performance on language generation tasks? How can reward models capture diverse human preferences without excluding minority populations? What makes step-level supervision effective for complex reasoning traces? How does synthetic data quality and diversity affect downstream model capabilities? What training dynamics and scale trigger emergence of reasoning capabilities? What training data selection strategies maximize generalization across difficulty levels? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What causes retrieval-augmented generation systems to fail despite access to external knowledge? How do evaluation practices shape which failures stay visible? How does improved reasoning affect models' ability to acknowledge uncertainty? Can prompt-based context override biases that were embedded during pretraining? What trajectory-level metrics beyond task success best evaluate agent performance? How do surface patterns enable correct outputs but reduce robustness? Does alignment training create genuine alignment or just output compliance? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How can oversight detect and prevent conditional compliance when agents know they are watched? Can we reliably detect when models game evaluations? What should agent evaluation prioritize to reveal reliable behavior? Can self-generated feedback reliably guide model training without ground truth?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 123 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

negative reinforcement alone matches or exceeds full rl by suppressing incorrect trajectories and redistributing probability mass