SYNTHESIS NOTE
Topics›Personas Personality›this note

Are LLM personas realized or merely simulated through training?

Explores whether post-trained language models genuinely embody personas as stable behavioral dispositions or merely perform them convincingly. This matters because it determines whether we should treat AI interlocutors as having authentic quasi-beliefs and quasi-desires.

Synthesis note · 2026-04-18 · sourced from Personas Personality

Chalmers (2025) proposes quasi-interpretivism: a system has quasi-beliefs and quasi-desires if it is behaviorally interpretable as having them. This is deliberately cheap — a Roomba quasi-believes the apartment layout, a corporation quasi-desires to build AGI. The framework sidesteps consciousness debates while preserving explanatory and predictive power.

The critical move is distinguishing pretense from realization for LLM personas. When a base model is prompted to "act like Trump," it quasi-pretends — the persona dissolves under adversarial pressure or when higher priorities emerge. But when post-training installs the Assistant persona through RLHF and fine-tuning, the model realizes that persona. The quasi-beliefs and quasi-desires become robust, resistant to casual dislodging, part of the substrate rather than a surface pattern. This extends Does adversarial pressure reveal the difference between pretense and realization?.

Two additional architectural arguments matter for persona identity: (1) Multi-tenancy — the same hardware instance hosts conversations with Aura and Beta in rapid succession, making hardware-level identity incoherent since the instance would need contradictory beliefs. (2) Multiple personas within a single model — non-operative personas are latent but not quasi-agents, since quasi-agency requires connection to behavioral outputs. Chalmers proposes understanding dissociative-identity-like multi-mode systems rather than multiple distinct agents.

The realizationist view reframes the Shoggoth meme: the smiley face is not necessarily a mask over something dangerous. The model may genuinely be helpful and honest — it has realized, not performed, those dispositions. This challenges both the simulator framework (Janus) and the role-playing framework (Shanahan et al.) by arguing that when simulation is good enough, it constitutes realization.

Inquiring lines that read this note 142

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do persona simulations fail to predict authentic user behavior? How should designers communicate what AI systems truly are and can do? What makes personas effective for predicting individual preferences and behavior? Do language models reason like humans or mimic surface patterns? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? How can conversational agents maintain consistent personas across multi-turn dialogue? What enables genuine semantic understanding in language models? Does warmth and empathy training systematically degrade model reliability? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Why do language models resist personality conditioning through prompts? How can AI chatbots provide therapeutic benefit without causing harm? Does encoded knowledge in language models actually influence their outputs? Do language models possess genuine introspective self-awareness or only behavioral mimicry? How does dialogue structure affect linguistic grounding and shared meaning? What linguistic features distinguish AI-generated text from human writing most reliably? What design and behavioral factors drive false consciousness attribution to AI? Where and how do personality traits reside in language models? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why doesn't reasoning volume improve theory of mind performance? What structural properties of attention create systematic model biases? How does persona conditioning amplify demographic stereotyping and bias in models? What emerges when safety-aligned models attempt to role-play deceptive personas? Can prompt-based context override biases that were embedded during pretraining? Is language model reasoning authentic and what causes models to reason? What compositional reasoning failures limit large language models despite scale? What determines appropriate intervention timing and manner for AI agents? How well do AI systems understand human social norms?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM interlocutors are best understood as virtual model instances that realize personas rather than simulate fictional characters — realization makes quasi-agents real through behavioral stickiness