SYNTHESIS NOTE
Topics›this note

Do large language models actually commit to a single character?

Explores whether LLMs pick and hold a fixed character or instead sample from multiple consistent possibilities. Tests reveal that regenerated responses differ while remaining consistent with context, challenging intuitive assumptions about how dialogue agents work.

Synthesis note · 2026-04-15 · sourced from Role-Play with Large Language Models

Shanahan constructs a simple but decisive behavioral test. Have an LLM-based dialogue agent play 20 questions — the agent "thinks of" an object and the user asks yes/no questions. After several rounds, ask the agent to reveal the object. It names something consistent with all previous answers. Now regenerate that response. The agent names a different object, also consistent with all previous answers.

This phenomenon is incompatible with any view that treats the agent as having committed to a specific object at the start of the game. A human playing 20 questions picks an object, holds it in mind, and answers questions from that fixed commitment. The LLM never picks. It maintains a set of objects consistent with the accumulated constraints — what Shanahan calls a superposition — and samples from that set at the moment of reveal. The same logic extends from objects to characters: the agent never commits to being a specific character with specific properties. It maintains a distribution over consistent characters and generates behavior sampled from that distribution.

The test is portable. Any feature that appears settled in one generation but changes on regeneration (while remaining consistent with context) is evidence of superposition rather than commitment. This has been observed in personality traits, stated preferences, claimed memories, and emotional dispositions of dialogue agents. The philosophical consequence is that attributing fixed psychological properties to an LLM conversation state is category-mistaken: the system has a distribution over properties, not a property. What appears stable is a high-probability region of the distribution, not a fact about an underlying entity.

Inquiring lines that read this note 142

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What articulatory and acoustic information does speech preserve that transcription destroys? What compositional reasoning failures limit large language models despite scale? Is language model reasoning authentic and what causes models to reason? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? What linguistic features distinguish AI-generated text from human writing most reliably? Why do persona simulations fail to predict authentic user behavior? Does encoded knowledge in language models actually influence their outputs? Can language models build genuine grounding through interaction? What mechanisms preserve shared understanding in evolving conversations? How can conversational agents maintain consistent personas across multi-turn dialogue? Why do LLM recommenders underperform collaborative filtering despite their capabilities? What prevents conversational agents from taking initiative in dialogue? Why do language models resist personality conditioning through prompts? Why do token-level mechanisms matter for learning to reason? Does preference optimization systematically degrade conversational grounding in language models? Why do some clarifying approaches produce understanding while others just satisfy? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? How does decomposing tasks improve reasoning and prevent failure propagation? Do language models respond to social pressure and face-saving like humans? Do language models learn genuine understanding or just surface patterns? Can diffusion models match autoregressive performance on language generation tasks? How do prompt design choices influence model reasoning and performance? How do multi-agent LLM systems fail distinctly compared to single agents? Why is dynamic grounding necessary for achieving true mutual understanding in dialogue? How does dialogue structure affect linguistic grounding and shared meaning? Can prompt-based context override biases that were embedded during pretraining? Do writers recognize when AI writing assistance alters their expressed stance? What makes personas effective for predicting individual preferences and behavior? Where and how do personality traits reside in language models? Why don't LLMs reliably translate capability into accurate outputs? Do language models reason like humans or mimic surface patterns? Can memory architectures handle ultra-long context better than attention? Does alignment training create genuine alignment or just output compliance? How does evaluation scope and dimensionality affect what we measure? Should agents decouple planning from perception grounding for better performance? What types of diversity prevent reasoning systems from collapsing? How do standardized protocols improve multi-agent coordination and reliability? Do language models lack essential therapeutic presence and engagement? How can AI chatbots provide therapeutic benefit without causing harm? How do spurious versus genuine rewards shape model reasoning and behavior?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 149 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the 20-questions regeneration test falsifies any committed-character view of LLM behavior