Do LLM judges systematically favor arguments from other LLMs?
When LLMs evaluate debates between LLM-generated and human arguments, do they show measurable preference for LLM-authored content? Understanding this bias matters because it affects every AI feedback loop used to train models.
When LLMs-as-judges were asked to score the same debates that human annotators scored, they picked the LLM as winner 62% of the time on average. Humans split 39% human / 37% LLM, with 24% draws. GPT-4o, the most accurate of the LLM judges, still picked the LLM 55% versus humans' 37% — and produced only 2% draws to humans' 24%. This is a same-kind-prefers-same-kind bias of substantial magnitude, layered on top of the four judge biases already catalogued elsewhere.
This is a tension because it bites every pipeline that uses LLMs to evaluate LLM output. Automated debate-quality scoring, RLHF-from-AI-feedback (RLAIF) loops, self-evaluation regimes, multi-agent debate frameworks that score each other's contributions — all inherit the bias. The result is a calibration ceiling: an evaluation pipeline whose output systematically over-credits LLM-authored arguments produces feedback signals that train models to produce more of what LLM judges over-credit, in a closed loop.
This sharpens Can LLM judges be fooled by fake credentials and formatting?. The four catalogued biases are exploitable by adversaries; same-author preference is a structural bias that needs no adversary. It activates whenever LLM-authored content is in the evaluation pool, which is to say, in every contemporary RLAIF pipeline.
It also bears on When does debate actually improve reasoning accuracy?. The Thin Line evidence shows the judge-side mechanism for that amplification: when contested-domain arguments are scored by LLM judges, the LLM-authored arguments win disproportionately, regardless of substantive merit. Multi-agent debate frameworks that close the loop with LLM judges are not just amplifying errors — they are amplifying their own preferred argument style.
The internal-consistency finding compounds the problem: humans' consistency between argument-strength scores and chosen winner was 73%; the LLM average was 55%. Even when the model assigned high strength scores to a human argument, it would often pick the LLM as winner anyway. The bias operates at the winner-selection step, downstream of component-level scoring.
Enrichment (2026-09-24, from Arxiv/RLVR): The vault now holds one measured RLAIF loop against a weak judge: in 2608.17776 the single-player baseline "quickly hacks the judge", and debate training avoided that on math (Can debate training prevent reward hacking by weaker judges?). Keep the two results apart. The 62% figure above is for judges scoring human-versus-LLM debates; the RL paper reports hacking of a frozen weaker judge and its excerpt gives no same-author analysis, so it neither confirms nor tests the preference described here. Whether an adversarial critic stays a guard or becomes a persuasion contest under a judge with this bias is the open part (Does debate prevent reward hacking without ground truth?), logged as debate training keeps a weak judge from being hacked on verifiable math while LLM judges favor LLM-authored arguments in contested domains — the protection may not survive where nothing can be checked.
For writing about evaluation infrastructure, the operational implication: evaluation by LLM judges of LLM output is not a substitute for human evaluation. Where LLM-as-judge pipelines are unavoidable, they need calibration corrections derived from human-labeled validation sets, applied per-task.
Inquiring lines that read this note 32
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why don't LLMs reliably translate capability into accurate outputs?- Where do LLMs succeed at generation but struggle with evaluation?
- Do LLMs match top human creative writers in literary quality?
- Why do LLMs excel at generation but struggle with evaluation?
- Can LLMs reliably assess the quality of ideas they generate?
- Do LLMs generate more novel ideas than they can evaluate?
- Why do LLM judges assign high argument strength scores yet pick LLM winners anyway?
- Does LLM judge preference for LLM arguments amplify errors in contested factual domains?
- How do calibration and reliability differ in LLM judge evaluations?
- Why do LLMs show gender bias but humans evaluators do not?
- Why do LLM judges show more extreme sycophancy bias than humans?
- What four exploitable biases make current LLM judges vulnerable to zero-shot attacks?
- Can LLM judges be trained to think more rigorously during evaluation?
- What other evaluation biases exist in LLM judge systems?
- What biases do single large LLM judges introduce into comparisons?
- What biases might an LLM judge introduce into an on-policy alignment process?
- What systematic biases do LLM judges introduce into AI-evaluated debates?
- What are the four catalogued biases that make LLM judges vulnerable to prompt attacks?
- How does LLM judge bias amplify errors in multi-agent debate on contested factual questions?
- Why is consistency between argument scoring and winner selection lower for LLM judges?
- Why do LLM judges systematically favor outputs from their own model family?
- Which biases in LLM judges are exploitable through presentation alone?
- Do smaller LLM judge panels outperform single large judges in practice?
- Can an LLM judge's bias be reduced through prompting or other interventions?
- Do LLM judges systematically favor arguments from other LLMs?
- Why do LLMs systematically prefer text from their own family?
- How do LLM judges' built-in biases influence the policies they help align?
- Can an LLM judge reliably report its own biases rather than remove them?
- Can LLMs evaluate logical argument quality in debates they themselves can win?
- How sensitive are LLM bias measurements to analysis choices?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
fifth bias to add, structural rather than adversarial
-
When does debate actually improve reasoning accuracy?
Multi-agent debate shows promise for reasoning tasks, but under what conditions does it help versus hurt? The research explores whether debate amplifies errors when evidence verification is missing.
judge-side mechanism for the amplification
-
Can debate training prevent reward hacking by weaker judges?
When an LLM policy trains against a weaker judge that supplies rewards, does the judge's systematic errors get exploited? This explores whether adversarial debate between generator and critic can sustain judge performance where single-player training fails.
a measured RLAIF loop against a frozen weaker judge; no same-author analysis in its excerpt
-
Does debate prevent reward hacking without ground truth?
Debate training reduced hacking in math tasks with verifiable answers, but the paper's own stated limit is whether this protection extends to domains where no correct answer exists to check against.
OPEN question: whether this bias turns an adversarial critic into a persuasion contest
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Humans or LLMs as the Judge? A Study on Judgement Biases
- The Thin Line Between Comprehension and Persuasion in LLMs
- Can Large Language Models Capture Human Annotator Disagreements?
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts
- Argument Collapse: LLMs Flatten Long-Form Public Debate
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
Original note title
LLMs-as-judges systematically prefer LLM-generated arguments over human ones — biasing any AI-evaluated debate pipeline