Analyzing and Correcting Benevolence Bias in Large Language Models

Paper · arXiv 2608.24912 · Published July 26, 2026
Role-Play and Persona Behavior

Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a “malicious persona” stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate.

Introduction. Large language models (LLMs) are increasingly used as stand-ins for human respondents across the social and behavioural sciences: to answer opinion polls, to simulate survey participants, to populate agent-based models of social processes, and to supply synthetic samples where collecting human data at scale is hard [1–4]. These uses promise faster, cheaper and larger studies, and they rest on one idea: that asking a model to answer as a given kind of person yields answers resembling what real people of that kind would say. Whether today’s aligned models live up to this idea, and on which questions, has never been tested systematically. Answering that question is what turns LLM-based simulation from a promising idea into a dependable method. The question matters because the same models are also built for a different job: from chat assistants to policy tools, being “helpful” and “harmless” is a core design goal.

Discussion / Conclusion. Across 18 widely used large language models and four major social-science datasets, we find a consistent shift of simulated human answers toward the benevolent end on value-laden questions: 83 of 108 BTB cells and 75 of 108 BWR cells are positive, concentrated on social desirability (mean BWR “ 0.565) and harm aversion (mean BWR “ 0.569), while emotional softening is close to null (mean BWR “ 0.494). Model rankings agree across the three general survey datasets (Spearman’s ρ between 0.54 and 0.63), so the benchmark captures a property of the models, not of any one dataset. We call this shift benevolence bias. It is not a thin layer of style on top of an otherwise faithful value representation, but a prior the model brings to every persona it adopts and every survey it answers; naming and measuring it is what makes it something researchers can plan around. Where the bias does and does not appear points to its source.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do embedding systems fail to capture task-relevant relationships? Why don't LLMs reliably translate capability into accurate outputs? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Does alignment training create genuine alignment or just output compliance? How can conversational agents maintain consistent personas across multi-turn dialogue? Do language models learn genuine understanding or just surface patterns? Do language models reason like humans or mimic surface patterns? How well do AI systems understand human social norms? Does preference optimization systematically degrade conversational grounding in language models? Why do persona simulations fail to predict authentic user behavior? Do language models lack essential therapeutic presence and engagement? Can AI systems distinguish genuine empathy from simulated emotion? How do agent-learned skills transfer and improve across different tasks? What drives appropriate trust calibration in personalized AI systems? Does warmth and empathy training systematically degrade model reliability?