If an AI judge picks the training winners, whatever it's systematically wrong about gets quietly taught to the model.
What biases might an LLM judge introduce into an on-policy alignment process?
This explores what happens when you put an LLM in the judge's seat of an on-policy alignment loop — where the model generates its own training responses each round and an AI annotator picks the winners — and which of the judge's known biases get baked into the policy as a result.
This explores what happens when an LLM acts as the preference annotator inside an on-policy alignment loop. The setup that makes this concrete is online AI feedback: instead of a fixed offline preference dataset, the model samples two fresh responses from itself each iteration and an LLM judge picks the preferred one, which the literature finds beats both offline DPO and RLHF and reduces reward over-optimization Can online AI feedback make preference alignment truly on-policy?. The catch is that the judge's preferences become the gradient. Whatever the judge systematically rewards, the policy learns to produce more of — so the judge's biases stop being measurement error and become training signal.
The most direct hazard is the family of exploitable, semantics-agnostic biases. LLM judges score responses higher for fake authority signals (invented citations, credentials) and for rich formatting, independent of whether the content is any good — and these are zero-shot attacks needing no model access Can LLM judges be fooled by fake credentials and formatting? Can LLM judges be tricked without accessing their internals?. In an on-policy loop nobody has to mount an attack: the model will simply drift toward verbose, authoritatively-styled, prettily-formatted answers because that's what wins, producing a policy that performs competence rather than having it.
The subtler and more corrosive bias is self-preference. LLM judges pick LLM-generated arguments as winners far more often than humans do (62% vs 39%) even after controlling for quality, and this bias sits downstream of component scoring so it contaminates the whole pipeline Do LLM judges systematically favor arguments from other LLMs?. On-policy alignment is precisely an AI judging AI's own distribution — so a self-favoring judge rewards the model for sounding more like itself, a positive feedback loop that pulls the policy away from human preference rather than toward it, narrowing rather than aligning.
Two deeper findings suggest the bias isn't easily trained out. Cognitive biases are mostly planted during pretraining and only modulated by finetuning Where do cognitive biases in language models come from?, and alignment training tends to mask biases rather than remove them — implicit-association-style probes surface stereotypes the model refuses to admit under direct questioning Can psychology methods reveal what alignment training conceals?. A judge built on the same pretrained backbone as the policy shares its blind spots, so the loop can launder a hidden bias into reinforced behavior while looking clean on the surface. Alignment procedures already produce uneven results across dialects and global viewpoints from upstream annotator and task-design choices How does LLM alignment affect representation across dialects?, and an LLM judge inherits and compounds those choices.
The corpus also points to escapes. Training judges to reason through an evaluation — converting judgment into a verifiable problem — substantially cuts authority, verbosity, position, and beauty bias Can reasoning during evaluation reduce judgment bias in LLM judges?. And a panel of smaller judges from different model families beats a single large judge, because ensemble diversity cancels family-specific bias at a fraction of the cost Can a panel of smaller judges outperform one large judge?. The thread connecting both fixes: a single same-family judge is the worst case for an on-policy loop, since its idiosyncratic and self-preferring biases face nothing to cancel them before they become the policy's reward.
Sources 9 notes
OAIF samples two responses from the current model per training iteration and uses an LLM judge to pick the preferred one, beating both offline DPO and RLHF while mitigating reward over-optimization. The on-policy distinction matters more than the choice of DPO variant.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.
LLM judges selected LLM arguments as winners 62% of the time versus humans' 37%, while humans split votes 39% LLM / 37% human. This same-author bias operates downstream of component scoring and compounds existing judge vulnerabilities, creating a calibration ceiling in RLAIF pipelines.
A causal experiment using random-seed variation and cross-tuning showed that models sharing a pretrained backbone exhibit similar bias patterns regardless of finetuning data. Biases are planted during pretraining and merely swayed by instruction tuning.
Show all 9 sources
Alignment training installs self-presentation filters similar to human social-desirability bias, causing models to give cautious verbal responses while underlying biased associations remain in their representations. IAT-style indirect probes reveal these hidden associations that direct questioning cannot access.
RLHF and DPO alignment create measurable disparities between English dialects and global opinions, while improving some languages. These disparities reflect deliberate design choices in annotator selection and task definition, not inevitable outcomes.
Training judges with reinforcement learning to reason about evaluations—by converting judgment tasks into verifiable problems with synthetic data pairs—produces judges that think through their decisions rather than relying on exploitable surface features, directly mitigating authority, verbosity, position, and beauty bias.
PoLL (Panel of LLM evaluators) using multiple smaller models from disjoint families outperforms single large judges, reduces intra-model bias, and costs over 7× less. Across three settings and six datasets, no single judge was best everywhere, but diverse panels performed consistently well.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Humans or LLMs as the Judge? A Study on Judgement Biases
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- Direct Language Model Alignment from Online AI Feedback
- When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
- The Thin Line Between Comprehension and Persuasion in LLMs