Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

Paper · arXiv 2609.16589 · Published September 15, 2026
Reinforcement Learning

As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes (“swing”), yet stubbornly ignore explicit instructions to correct ingrained biases (“rigidity”). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core.

Introduction. First, Do LLMs have values? To answer this, we evaluate LLMs using largescale socio-psychological surveys. However, we recognize that LLM values cannot be accurately described by a single static vector. LLMs exhibit “swing” behavior under repeated queries, and their value expression is fundamentally a distribution, not a fixed point. Therefore, we shift the evaluation paradigm from a single individual to a macrolevel population [12, 13]. We administer 150,000 queries to each of 106 LLMs across 625 designed scenarios, modeling each model’s value expression as a full statistical distribution. To evaluate LLM values in a human-interpretable way, we ground our approach in two established social science frameworks: the World Values Survey (WVS) [14, 15] and Schwartz Value Theory [16, 17]. We use data from Wave 7 of the WVS (2017- 2022; hereafter WVS-7). Wave 7 is the most recent, largest, and most geographically comprehensive wave to date, comprising approximately 95,000 valid respondents.

Discussion / Conclusion. Traditional safety benchmarks typically evaluate LLMs using single-turn, static questionnaires, implicitly treating the model as a single human subject with a fixed persona 3.2 LLMs as Aligned Tools Rather Than Human Surrogates Ultimately, this study reframes LLM values not as fixed, human-like personalities, but as dynamic, probabilistic fields. We reveal that value crystallization is a deterministic byproduct of parameter compression, which, when heavily constrained by safety protocols, fundamentally disqualifies LLMs as surrogates for diverse human populations. Furthermore, by uncovering the dual-process dynamics of contextual framing and cognitive reasoning during inference, we transition value alignment from opaque trial-and-error into a mechanistic science. This structural understanding is crucial: it empowers us to reduce the “alignment tax” through precise, prescriptive interventions, ensuring that generative models remain highly controllable, transparent tools rather than unexamined uninterpretable black boxes.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can local safety checks guarantee system-level behavioral safety? Do language models reason like humans or mimic surface patterns? Is language model reasoning authentic and what causes models to reason? Why don't LLMs reliably translate capability into accurate outputs? Should agents decouple planning from perception grounding for better performance? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Does alignment training create genuine alignment or just output compliance? Does preference optimization systematically degrade conversational grounding in language models? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? What emerges when safety-aligned models attempt to role-play deceptive personas? How well do AI systems understand human social norms? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Where and how do personality traits reside in language models?