SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Can models trained on many imperfect experts outperform everyone?

Can generative models trained on diverse, biased experts achieve better performance than any individual contributor? This explores whether aggregating diverse perspectives during training acts as implicit denoising.

Synthesis note · 2026-02-22 · sourced from Training Fine Tuning

The Transcendence paper formalizes a surprising property: generative models trained on many experts with diverse capacities and biases can outperform any single expert. The mechanism is implicit majority voting. When trained on diverse human players (chess), the model's cross-entropy optimization converges on the consensus behavior — which, by the wisdom-of-the-crowd effect, is often better than any individual contributor.

Low-temperature sampling is the key enabler. At low temperature, the model's output distribution concentrates on its highest-probability predictions — the consensus. This is formally equivalent to a majority vote. The advantage is primarily due to performing much better on a small subset of states — likely the critical, outcome-determining positions where individual human biases diverge most and the crowd wisdom is most valuable.

Diversity in the training data is a necessary condition. Without diversity, there is no denoising — a model trained on clones of one expert can only approach that expert's level. The practical conditions for transcendence: (1) diverse training sources with different biases, (2) a task where individual biases are uncorrelated (so they cancel under aggregation), and (3) low-temperature decoding to extract the consensus.

This connects to but is distinct from Why does majority voting outperform more complex inference methods?. That note describes inference-time majority voting over multiple samples from one model. Transcendence describes training-time majority voting implicitly encoded in a single model's weights through diverse training data. The mechanism is analogous — aggregation denoises — but operates at different timescales.

The implication for LLM training is provocative: the "average" of many imperfect human demonstrations may be better than any individual human demonstration, provided the imperfections are diverse rather than correlated. This challenges the assumption that training data quality should be maximized per-example; quantity and diversity of perspectives may matter as much as individual quality.

Inquiring lines that read this note 31

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? When do multi-agent systems outperform single frontier models? How well do AI systems understand human social norms? How does evaluation scope and dimensionality affect what we measure? How does synthetic data quality and diversity affect downstream model capabilities? What types of diversity prevent reasoning systems from collapsing? What do systematic disagreements between annotators reveal about ground truth? What training dynamics and scale trigger emergence of reasoning capabilities? Can self-generated feedback reliably guide model training without ground truth? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How should designers communicate what AI systems truly are and can do? How do social dynamics distort aggregated online ratings? What capability trade-offs arise from domain specialization through fine-tuning? How do surface patterns enable correct outputs but reduce robustness? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How can we prevent synthetic data from contaminating statistical inference and corpora? How can conversational agents maintain consistent personas across multi-turn dialogue?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 212 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

generative models transcend their training experts through implicit majority voting that denoises diverse human biases