Can models trained on many imperfect experts outperform everyone?
Can generative models trained on diverse, biased experts achieve better performance than any individual contributor? This explores whether aggregating diverse perspectives during training acts as implicit denoising.
The Transcendence paper formalizes a surprising property: generative models trained on many experts with diverse capacities and biases can outperform any single expert. The mechanism is implicit majority voting. When trained on diverse human players (chess), the model's cross-entropy optimization converges on the consensus behavior — which, by the wisdom-of-the-crowd effect, is often better than any individual contributor.
Low-temperature sampling is the key enabler. At low temperature, the model's output distribution concentrates on its highest-probability predictions — the consensus. This is formally equivalent to a majority vote. The advantage is primarily due to performing much better on a small subset of states — likely the critical, outcome-determining positions where individual human biases diverge most and the crowd wisdom is most valuable.
Diversity in the training data is a necessary condition. Without diversity, there is no denoising — a model trained on clones of one expert can only approach that expert's level. The practical conditions for transcendence: (1) diverse training sources with different biases, (2) a task where individual biases are uncorrelated (so they cancel under aggregation), and (3) low-temperature decoding to extract the consensus.
This connects to but is distinct from Why does majority voting outperform more complex inference methods?. That note describes inference-time majority voting over multiple samples from one model. Transcendence describes training-time majority voting implicitly encoded in a single model's weights through diverse training data. The mechanism is analogous — aggregation denoises — but operates at different timescales.
The implication for LLM training is provocative: the "average" of many imperfect human demonstrations may be better than any individual human demonstration, provided the imperfections are diverse rather than correlated. This challenges the assumption that training data quality should be maximized per-example; quantity and diversity of perspectives may matter as much as individual quality.
Inquiring lines that read this note 31
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking?- How does same-author bias interact with the four adversarial judge biases already documented?
- When do aggregated imperfect demonstrations fail to outperform the best expert?
- How do ensemble methods reduce bias in automated evaluation?
- Does disjoint family diversity actually cancel model-specific bias in evaluation?
- Can diverse expert demonstrations exceed the knowledge of any single expert?
- Can individually accurate agents still fail at population-level representation?
- Can small directional biases add up to meaningful population effects?
- How do experts select which other experts to trust?
- What happens when majority voting converges to a single answer?
- Why does low temperature sampling extract consensus from diverse training data?
- What conditions make training diversity better than individual expert quality?
- Does debiasing training data actually solve the bias problem in machine learning?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Can synthetic data diversity preserve the transcendence effect or does it collapse?
- How does mutual shaping through diverse training compare to population-level diversity effects?
- How does majority voting fail when reasoning samples lack genuine diversity?
- Can population-level distributions shift usefully even when individual prediction fails?
- Do draft-and-revise loops work better when guided by unresolved constraints than by diffusion-style denoising?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why does majority voting outperform more complex inference methods?
Simple majority voting across independent samples often matches or beats sophisticated alternatives like Best-of-N and sequential revision. What makes this basic approach so hard to beat for reasoning models?
inference-time voting analog; this is the training-time version
-
Does voting discard useful reasoning from losing chains?
When multiple reasoning chains compete through majority voting, intermediate steps from non-winning chains are discarded. Could extracting and mixing those intermediate facts improve both the final answer and our ability to understand the reasoning?
shows limits of pure voting; transcendence may have similar limits
-
Does training on AI-generated content permanently degrade model quality?
When generative models train on outputs from previous models, do the resulting models lose rare patterns permanently? The question matters because future training data will inevitably contain synthetic content.
counterpoint: while diversity enables transcendence, synthetic data collapses diversity
-
Can generative and discriminative models reach agreement?
Generative and discriminative decoding often produce conflicting answers. Can a game-theoretic framework force these two complementary procedures to reconcile their predictions into a single, more reliable output?
related consensus mechanism: transcendence achieves consensus across diverse training experts at training time, while Consensus Game achieves consensus between generative and discriminative decoding modes at inference time; both extract a signal more reliable than any single perspective
-
Can a quorum of validators really provide independent judgment?
If multiple validators share training data, prompts, evidence sources, or infrastructure, their agreement may reflect shared causes rather than independent confirmation. This could make quorum-based systems less reliable than they appear.
the deployment-time inverse of the uncorrelated-biases condition: validators chosen at inference time may share weights, prompts, retrieval and evidence, so a vote among them need not denoise anything; that note reports no measured correlation
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Transcendence: Generative Models Can Outperform The Experts That Train Them
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- Multistep Consistency Models
Original note title
generative models transcend their training experts through implicit majority voting that denoises diverse human biases