The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50% of personas fail to express their assigned WVS profiles from the outset, while 2-7% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges.
Introduction. There is a growing interest in using LLMs as stand-ins for human populations. Prior work has used LLMs to simulate individuals with specific demographic attributes or personality traits (Argyle et al. 2023; Tjuatja et al. 2024; Wang et al. 2025), as well as collective phenomena such as group dynamics and organizational behavior (Park et al. 2024; Zhou et al. 2024). In these settings, LLMs are increasingly positioned as scalable proxies for human judgment, opinion, and behavior (Choi, Young, and Ferrara 2026; Tjuatja et al. 2024; Park et al. 2024). However, real-world societies are shaped not only by demographic diversity but also by diverse and often conflicting value systems(Schwartz 1992). Therefore, the validity of such simulations depends on whether LLM agents can reliably instantiate and sustain culturally grounded value profiles during interaction.
Discussion / Conclusion. While LLM-based simulations increasingly model individuals and populations, our results show that conversational plausibility can mask failures of representational fidelity. They may therefore reproduce long-standing external validity problems, including overgeneralization from WEIRD populations (Henrich, Heine, and Norenzayan 2010). These risks might get amplified in synthetic populations, digital twins, and policy simulations, where unfaithful value profiles can distort estimates of public opinion, intervention effects, or group behavior. Such distortions may especially affect marginalized or culturally underrepresented groups and contribute to algorithmic monoculture, where downstream studies inherit the same model-specific biases (Kleinberg and Raghavan 2021; Saha and Choudhury 2025; Pandey, Saha, and Choudhury 2025; Saha, Pandey, and Choudhury 2025; Saha et al. 2025). Also, since richer demographic prompting is not a sufficient safeguard, LLM-based simulations should be treated as virtual respondents, but as modelmediated instruments requiring validation against the populations they claim to represent (Bisbee et al. 2024; Ziems et al. 2024; Madden 2025).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do embedding systems fail to capture task-relevant relationships? Do language models learn genuine understanding or just surface patterns? Do language models reason like humans or mimic surface patterns?- How do value distributions differ across model families and training scales?
- What structural coherence exists in LLM preference systems and value hierarchies?
- Why does weakening communication fail but weakening belief succeeds?
- Should safety constraints trade off against representing authentic human value diversity?
- Should simulated value-based discussions be validated against real human populations?
- How does face-saving behavior let AI mimic community participation without joining it?
- What role does human response variation play in LLM simulation accuracy?
- Do emotion-driven actions in agent simulators capture genuine belief revision or just reactive behavior?
- Can agent-based simulators replace real-user A/B testing for studying recommendation system harms?
- Can controllable latent variables in simulators ground them to realistic conversation?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Why do longer forecasting horizons degrade LLM accuracy in role-play?
- How do multi-agent LLM systems fail at coordination and role consistency?
- Why does silent agreement occur so often in multi-agent LLM systems?
- Can parallel agents or complementary mechanisms replace single-human interrogation of LLMs?