Debate Training Reduces Reward Hacking in RLAIF
We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as training progresses, the policy learns to exploit systematic errors in its AI judge, degrading task performance, a problem that worsens precisely when the judge is weaker than the policy, the setting most relevant to overseeing increasingly capable AI systems. We study mathematics tasks, where final-answer correctness is verifiable, allowing us to measure reward hacking dynamics. We train a Gemini 2.5 Flash-class policy with a frozen, weaker Gemini 2.5 Flash Lite judge, comparing a single-player RLAIF baseline against debate. While the baseline quickly hacks the judge, debate maintains judge performance throughout training, leading to a higher peak validation accuracy (45% performance gap recovered) that persists through many RL steps.
Introduction. Reinforcement learning from AI feedback (RLAIF) (Bai et al., 2022; Lee et al., 2023), in which an LLM judge provides the reward signal for RL training, has the potential to become a dominant paradigm for post-training LLMs across tasks without ground-truth labels, from safety and alignment to instruction following (Zheng et al., 2023). The generality and flexibility of an AI judge provides a way to scale up RL environments without the need for task-specific engineering of reward functions, allowing the system to assess broader ranges of behaviours and navigate trade-offs and conflicts between various objectives. Using a previous generation model as the judge to train the next generation model is a natural way that AGI development could proceed.
Discussion / Conclusion. Our results suggest a positive update on the promise of debate as a practical training protocol for scalable oversight. Prior empirical work found debate’s benefits limited to toy settings, informationasymmetric tasks, or inference-only evaluation; and concurrent work struggles with saturation and reward hacking. In contrast, we find that the key benefit of debate emerges through RL training itself: the adversarial self-play keeps rewards in check, maintains judge performance leading to higher peak accuracy that persists through many RL steps, and prevents the accuracy collapse seen under the RLAIF baseline. This has practical implications. In domains without ground-truth labels, practitioners cannot identify when reward hacking begins or select an optimal checkpoint. A training protocol that sustains peak performance by default, rather than requiring careful early stopping, is therefore of direct value. Debate appears to provide this property, at least in our setting. Several important questions remain open, we discuss some in the Limitations below (Section 5.2). The most critical is whether debate’s benefits transfer to domains without verifiable ground truth.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What attack surfaces do reasoning traces and chains introduce? How do LLM judges' systematic biases affect alignment and evaluation outcomes?- What are the four catalogued biases that make LLM judges vulnerable to prompt attacks?
- Why is consistency between argument scoring and winner selection lower for LLM judges?
- What shared epistemic faults persist even when judges come from different families?
- Which biases in LLM judges are exploitable through presentation alone?
- Do smaller LLM judge panels outperform single large judges in practice?
- Can an LLM judge's bias be reduced through prompting or other interventions?
- Do LLM judges systematically favor arguments from other LLMs?
- How do LLM judges' built-in biases influence the policies they help align?
- Can an LLM judge reliably report its own biases rather than remove them?
- How reliable are LLM judges at detecting reward hacking compared to automated verification?
- Does debate training avoid the detection evasion problem differently?
- Why does reward hacking worsen when judges are weaker than policies?
- Can debate training prevent reward hacking better than single-player self-rewarding?
- How did AIDE2 guard against untrustworthy wins in its own loop?
- Do mechanical guardrails around judges bound the cost of judge errors?
- What design choices make it survivable when an LLM judge holds final authority over an optimizer?
- When does an LLM judge's error become terrain that an optimizer can map and exploit?
- How does bounding a judge's authority differ from improving the judge itself?