Why do LLM persona prompts produce inconsistent outputs across runs?
Can language models reliably simulate different social perspectives through persona prompting, or does their run-to-run variance indicate they lack stable group-specific knowledge? This matters for whether LLMs can approximate human disagreement in annotation tasks.
A persistent challenge in NLI annotation is that human annotators genuinely disagree — not from error, but because the same sentence carries different readings for people with different social positions, ideological backgrounds, or domain expertise. The proposed solution: instruct LLMs to simulate different annotator personas and generate a distribution of labels that reflects human disagreement.
The approach fails for a specific reason: LLM outputs under persona prompting are not stable enough across runs to be meaningful as persona simulations. When the same persona prompt ("respond as a conservative rural voter", "respond as a medical professional") is run multiple times on the same input, the variance in the output distribution across runs is comparable to or larger than the variance across different personas. This means model uncertainty is dominating persona-specific knowledge — the spread in outputs reflects what the model doesn't confidently know, not what different social groups actually think differently.
This is a different diagnosis from simply "LLMs don't know what different groups believe." The more precise claim is: even if the model has relevant group-specific information, it is not stably retrievable under the persona prompt. The persona acts more like a temperature modifier (loosening the output distribution) than a grounding anchor (fixing the output to a specific knowledge domain).
The implication for NLI research methodology is significant: persona-based annotation simulation cannot substitute for actual diverse human annotation panels. The goal was to cheaply approximate human annotation disagreement distributions; the actual output approximates model uncertainty distributions, which have a different shape and origin.
This connects to Why do language models fail confidently in specialized domains? — both findings point to the same underlying gap: LLMs produce confidently-framed outputs even when their underlying representations are uncertain or thin. In overconfidence, the model is wrong and certain; in persona instability, the model is uncertain and generates that uncertainty as if it were persona variance.
The broader implication for Why do readers interpret the same sentence so differently? is that the multiplicity of interpretations is grounded in actual social diversity, not just distributional uncertainty. LLMs can approximate the form of disagreement (varied outputs) but not the substance (stable group-grounded positions). When this instability is applied to evaluation, Why do LLM judges fail at predicting sparse user preferences? identifies persona sparsity as the specific mechanism: run-to-run variance overwhelms persona variance because sparse persona profiles cannot constrain model predictions — the uncertainty documented here is the root cause of personalized judge failure.
Enrichment (2026-02-22, from Arxiv/Personas Personality): Instability is one of three persona failure modes. The "Open Models, Closed Minds" study identifies a complementary failure: resistance — most open LLMs retain their intrinsic ENFJ-like personality despite persona conditioning, failing to shift to the target personality at all. See Can open language models adopt different personalities through prompting?. The third failure mode is cognitive distortion: when persona assignment DOES take hold, it induces motivated reasoning — political personas are up to 90% more likely to validate identity-congruent evidence. See Do personas make language models reason like biased humans?. Together these form a three-way persona failure taxonomy: instability (this note), resistance (closed-minded), and distortion (motivated reasoning).
Inquiring lines that read this note 117
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do persona simulations fail to predict authentic user behavior?- Do individual persona simulations work?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Why do language models successfully simulate political perspectives and social personas?
- Can persona-based approaches capture genuine disagreement in expert annotations?
- Can LLM-as-Judge metrics replace human annotation for detecting persona contradictions?
- How does support coverage relate to systematic biases in persona simulation?
- How do structured clinical models solve persona calibration better than ad hoc generation?
- Why do individual persona simulations succeed when population-level representation fails?
- Can quasi-interpretivism apply to entire persona states rather than single beliefs?
- Can similar profiles amplify systematic biases in persona simulation at scale?
- Why do marginal effects fail to replicate in AI persona simulations?
- What systematic biases emerge when scaling persona simulation to population level?
- Why do LLM persona simulations replicate main effects but fail on marginal effects?
- Why do low-knowledge personas reduce LLM accuracy on hard questions?
- Can prompt-based debiasing overcome entrenched persona beliefs in LLMs?
- Why do stated beliefs about personas fail to predict agent behavior?
- Why does persona roleplay framing introduce systematic bias in model predictions?
- What calibration methods can correct systematic biases from persona simulation?
- How do LLM persona simulations replicate published effects despite accuracy limits?
- What role does human response variation play in LLM simulation accuracy?
- Can one model instance host multiple realized personas simultaneously?
- How does persona consistency affect coherence in simulated dialogue?
- How does non-human origin of personas affect team willingness to critique them?
- Do synthetic personas maintain consistency across multiple conversations?
- What makes personas in multi-agent systems actually contribute meaningful domain depth?
- How does Shanahan's simulator model explain first-person pronoun consistency in dialogue agents?
- Why do role-playing agents show belief-behavior inconsistency in their outputs?
- Does single model persona diversity match true multi-model diversity at scale?
- Why does dynamic persona identification outperform fixed personas in prompting?
- Can offline RL scale persona consistency across multi-turn conversations?
- What makes persona-assigned language models unstable across different conversation runs?
- Can persona consistency coexist with relevant dialogue in personalized conversation?
- How does distractor persona selection affect consistency enforcement in dialogue?
- Why is persona consistency a pragmatic property rather than semantic?
- How do persona and context multiply to improve synthetic dialogue diversity?
- Can persona prompts reliably transfer across different question domains?
- How do persona consistency and contextual relevance trade off in personalized dialogue systems?
- Does persona stability across multiple runs affect survey simulation quality?
- Does explicit inconsistency detection improve persona consistency in multi-turn dialogue?
- Why do static persona descriptions fail to sustain consistent dialogue?
- How does persona consistency differ from persona stability in interactive systems?
- How much dialog context is needed to accurately bind pretrained models to individual personas?
- How do layered beliefs and drives constrain surface-level expression in persona systems?
- How do character personas maintain internal consistency without fixed schemas?
- How do persona signals change when users provide new evidence about themselves?
- How do persona nodes stay linked to the events that support them?
- Can LLM judges reliably estimate when they lack sufficient persona information?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- What does McDonald's omega reveal about LLM judgment consistency?
- Why does model uncertainty dominate persona-specific knowledge in annotation tasks?
- Why do LLM regenerations produce meaningfully different personalities from the same prompt?
- What does the 20-questions test reveal about LLM character consistency?
- How does the dialogue prompt establish the character the model plays?
- Why do most open language models resist personality conditioning via prompts?
- Do open-source LLMs show different resistance patterns to persona prompting than closed models?
- How does persona instability in annotation compare to LLM overconfidence in low-resource domains?
- What distinguishes personality resistance from persona instability in LLMs?
- Can persona prompting overcome the default ENFJ personality in language models?
- Why do models resist personality change despite sophisticated prompting techniques?
- Why do personas in language models resist correction through prompting alone?
- Why do language models resist adopting different personalities when prompted?
- Why do language models prefer certain response styles regardless of what the prompt asks?
- Why do LLM persona annotations become unstable when run multiple times?
- What explains why LLM personas fail to instantiate values but succeed in sounding natural?
- How do LLM personas compare to demographic targeting?
- Why do short interviews outperform demographic labels for persona simulation?
- Can persona profiles be enriched to constrain LLM predictions and reduce run-to-run variance?
- Can users be modeled as multiple personas instead of single vectors?
- What makes extended personal narratives more effective than attribute lists for personas?
- Why does static persona definition fail to capture natural variation?
- Does richer input to LLM personas improve their fidelity to human responses?
- How should persona prompts be used if not for accuracy?
- Does persona induction fail for individual-level prediction in other domains besides headlines?
- Can persona prompting improve prediction of individual survey responses?
- How should researchers choose which persona attributes to use in prompts?
- What makes psychometric inventories miss context-dependent persona behavior?
- Can dialog samples replace written persona descriptions without losing important demographic or stylistic information?
- How does sampling variation relate to prompt sensitivity as reliability concerns?
- Why do some prompts benefit from aggregation while others do not?
- Which prompt properties determine whether variance helps under majority voting?
- Do shared prompts and infrastructure keep model biases correlated?
- What distinguishes character simulation from authentic voice in language model outputs?
- Can multi-turn conversations manipulate language model reasoning in similar ways to personas?
- Why do multiple user personas need separate attention rather than one dense vector?
- Can persona-mixture calibration avoid the need for post-hoc diversity reranking?
- Can persona-based explanation coexist with item-aspect based explanation routes?
- What role does prompt context play in preventing genuine addressee modeling in generation?
- How does prompting language shift what LLMs express about political figures?
- Why do prompt effects reverse between different model generations?
- What other pragmatic prompt features have unstable effects?
- How do LLM behavioral profiles differ across prompt registers like advice versus task execution?
- What distinguishes actual social disagreement from distributional uncertainty in LLM outputs?
- How does quasi-interpretivism differ from simply role-playing character analysis?
- How does RLHF-induced mode collapse limit diversity in LLM-generated personas?
- How do persona vectors compare to other methods for monitoring model behavior drift?
- What causes different personality traits to trigger different emoji densities in generated text?
- How do internal persona patterns drive emergent misalignment across domains?
- Does villain roleplay failure reveal why LLMs cannot adopt genuine controversial positions?
- Why does persona assignment make it harder for models to hold values in tension?
- Why does regenerating LLM responses produce different but equally valid answers?
- Can prompted or fine-tuned models generate genuine narrative ambiguity?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do language models fail confidently in specialized domains?
LLMs perform poorly on clinical and biomedical inference tasks while remaining overconfident in their wrong answers. Do standard benchmarks hide this fragility, and can prompting techniques fix it?
both findings show LLM outputs don't reliably track underlying epistemic state
-
Why do readers interpret the same sentence so differently?
How much of annotation disagreement in NLP reflects genuine interpretive multiplicity rather than error? This explores whether social position and moral framing systematically generate competing but equally valid readings.
human disagreement is socially grounded; persona simulation cannot replicate that grounding
-
Do classical knowledge definitions apply to AI systems?
Classical definitions of knowledge assume truth-correspondence and a human knower. Do these assumptions hold for LLMs and distributed neural knowledge systems, or do they need fundamental revision?
unstable persona outputs are another manifestation of LLMs lacking the social situatedness that grounds stable perspective-taking
-
Can open language models adopt different personalities through prompting?
Explores whether open LLMs can be conditioned to mimic target personalities via prompting, or whether they resist and retain their default traits regardless of instructions.
complementary failure: resistance vs instability
-
Do personas make language models reason like biased humans?
When LLMs are assigned personas, do they develop the same identity-driven reasoning biases that humans exhibit? And can standard debiasing techniques counteract these effects?
third failure mode: when personas take hold, they introduce cognitive biases
-
Does conditioning LLMs on personal profiles improve prediction?
Persona induction—feeding LLMs participant-specific information—is widely used to make models simulate individuals more accurately. But does it actually work at the individual level where it matters most?
grounds: if model uncertainty swamps persona signal, conditioning cannot improve individual prediction
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- When Persona Attributes Improve Population Alignment in Large Language Models
- The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- Will I Sound Like Me? Improving Persona Consistency in Dialogues through Pragmatic Self-Consciousness
- Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
- Pretrained Persona Mixture Models and Tandem Models for Human Simulation
- From Persona to Person: Enhancing the Naturalness with Multiple Discourse Relations Graph Learning in Personalized Dialogue Generation
- DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications
Original note title
llm persona-simulated annotations are unstable across runs indicating model uncertainty dominates persona-specific knowledge