On Epistemic Diversity in Large Language Models

Paper · arXiv 2609.04835 · Published September 4, 2026
LLM Evaluations and Benchmarks

Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A system may give a correct answer while still narrowing users’ access to alternative valid answers, explanations, or reasoning routes. Drawing on the broader notion of epistemic diversity in philosophy and social epistemology, we formalize it in the context of LLMs as the range of valid answers, explanations, and reasoning routes that an LLM exposes to users. We argue that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks. We propose a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs, and operationalize it in two domains. We find that frontier LLMs often exhibit epistemic narrowness, repeatedly collapsing large valid answer spaces onto small canonical subsets. These findings suggest that LLM evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as an important dimension of model capability.

Introduction. Large language models (LLMs) are increasingly used not only as predictive tools, but as knowledge tools for answering questions, explaining concepts, drafting arguments, and supporting inquiry (Chatterji et al., 2025). In these settings, what matters is not only whether a model can produce a correct or acceptable answer. What also matters is the kinds of answers, explanations, and reasoning the model makes available to users. A system may be accurate and yet still be epistemically narrow if, in contexts where alternatives would be useful, it repeatedly presents only a few canonical routes despite the existence of multiple valid answers. Much of the existing literature on diversity in AI asks whether different demographic, cultural, or political groups are represented or treated fairly (Hardt et al., 2016; Barocas et al., 2023; Guo & Caliskan, 2021; Wang et al., 2025). Those are important concerns.

Discussion / Conclusion. We introduced epistemic diversity as a distinct dimension of language model evaluation: the range of valid answers, explanations, examples, concepts, and reasoning strategies that a model makes available to users. Unlike group-based diversity, epistemic diversity concerns coverage over valid answer spaces or answer classes. We formalized this idea through a framework that asks what makes an answer valid, why multiple valid answers arise, and how diversity should be measured under different interaction protocols. We operationalize this framework across ten models and two datasets, finding that frontier LLMs often exhibit epistemic narrowness, even when many valid alternatives exist. In the professions domain, models repeatedly concentrate on a small set of canonical individuals; in the proofs domain, they often return the same proof strategy despite the existence of accepted alternatives. These results suggest that models do not merely answer questions, but shape which knowledge becomes salient.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do embedding systems fail to capture task-relevant relationships? Do language models learn genuine understanding or just surface patterns? Do language models reason like humans or mimic surface patterns? How well do AI systems understand human social norms? Does preference optimization systematically degrade conversational grounding in language models? Why do persona simulations fail to predict authentic user behavior? How can we prevent synthetic data from contaminating statistical inference and corpora? How should designers communicate what AI systems truly are and can do? What compositional reasoning failures limit large language models despite scale? How does AI-generated content undermine authentic engagement on social platforms? What factors drive AI persuasiveness and how can it be mitigated? Does alignment training create genuine alignment or just output compliance? How does dialogue structure affect linguistic grounding and shared meaning? What types of diversity prevent reasoning systems from collapsing? How do surface patterns enable correct outputs but reduce robustness? How do LLM judges' systematic biases affect alignment and evaluation outcomes? When do multi-agent systems outperform single frontier models? Can multi-agent systems avoid converging on false agreement without deliberation?