Does an LLM commit to a single character or maintain many?
Explores whether language models lock into one personality or instead hold multiple consistent characters in a probability distribution that narrows over time. Matters because it changes how we interpret apparent inconsistencies in model behavior.
The simple role-play metaphor — one actor, one part — is too rigid for what LLMs actually do. Shanahan refines it using Janus's simulator framing: the LLM is a non-deterministic simulator capable of generating an infinity of characters (simulacra), and at any point during a conversation it maintains a superposition of simulacra consistent with the preceding context. The superposition narrows as the conversation proceeds: each new turn rules out characters inconsistent with what has been said, concentrating probability on an ever-smaller set.
The distributional view is more than a refinement — it changes the ontological picture. Under simple role-play, there is one character the system is playing, and the question is what that character's properties are. Under the superposition view, there is no single character until the conversation has proceeded far enough to collapse the distribution to near-determinacy. The system is simultaneously consistent with many characters, and the character that appears in any particular generation is a sample from the current distribution, not a reveal of a committed identity.
This explains observable phenomena that the single-character view cannot. When a user regenerates the model's output, the second generation may present a meaningfully different personality, stance, or knowledge state — while remaining consistent with the conversation so far. The system did not change its mind; it sampled a different point from the distribution. The 20-questions test formalizes this: the agent never "thought of" an object; it maintained a set of objects consistent with prior answers and generated one on the fly at the reveal, and will generate a different consistent one if asked again.
Inquiring lines that read this note 57
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What mechanisms preserve shared understanding in evolving conversations?- Why do LLMs fabricate continuity when users shift conversational frames?
- Can the same conversation coherently continue across different model versions?
- How does psychological continuity theory apply to identity across LLM conversation threads?
- Can distributional views explain when an LLM appears to change its mind?
- What property must remain constant to individuate an LLM across infrastructure changes?
- How does maintaining a superposition differ from committing to a character?
- Why do LLMs succeed at social roles without a stable self?
- How can multiple conflicting values coexist in a single LLM system?
- Why do LLM stories over-explain themes and favor single-track plots?
- Can one model instance host multiple realized personas simultaneously?
- Why is persona consistency a pragmatic property rather than semantic?
- How do character personas maintain internal consistency without fixed schemas?
- Why do LLM regenerations produce meaningfully different personalities from the same prompt?
- What does the 20-questions test reveal about LLM character consistency?
- How does the dialogue prompt establish the character the model plays?
- How does persona instability in annotation compare to LLM overconfidence in low-resource domains?
- Why do some open models resist personality conditioning while others don't?
- What distinguishes personality resistance from persona instability in LLMs?
- Why do models resist personality change despite sophisticated prompting techniques?
- Why do language models resist adopting different personalities when prompted?
- Why do LLM persona annotations become unstable when run multiple times?
- What explains why LLM personas fail to instantiate values but succeed in sounding natural?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- What does McDonald's omega reveal about LLM judgment consistency?
- How do alignment constraints affect whether LLMs show emotional flexibility?
- Does alignment training intensity push LLM personas from pretense toward realization?
- How does model capability relate to personality conditioning flexibility?
- How does RLHF-induced mode collapse limit diversity in LLM-generated personas?
- Do personality traits occupy consistent geometric structures across different LLM architectures?
- What causes different personality traits to trigger different emoji densities in generated text?
- How does semantic entanglement interact with personality dimension shifts during finetuning?
- Can we detect superposition in LLM personality traits and stated preferences?
- How do personality and language proficiency moderate the impact of linguistic alignment?
- Do open language models default to a single shared personality type?
- How does language condition affect model psychological profile consistency?
- How does embodiment relate to whether something can have a persistent identity?
- How does model weight freezing across users affect virtual instance individuation?
- Why do aligned models struggle with deceptive character traits more than cruelty?
- Why does persona assignment make it harder for models to hold values in tension?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do large language models actually commit to a single character?
Explores whether LLMs pick and hold a fixed character or instead sample from multiple consistent possibilities. Tests reveal that regenerated responses differ while remaining consistent with context, challenging intuitive assumptions about how dialogue agents work.
the empirical demonstration of superposition
-
Should we treat dialogue agents as role-playing characters?
Does the role-play framing successfully avoid anthropomorphism while preserving folk-psychological vocabulary for describing LLM behavior? This matters because it shapes whether we attribute genuine mental states to dialogue systems.
the simple role-play view this refines
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Role-Play with Large Language Models
- Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
- EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
- Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- What we talk to when we talk to language models
- PersLLM: A Personified Training Approach for Large Language Models
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
Original note title
an LLM is a non-deterministic simulator that maintains a superposition of simulacra rather than committing to a single character