Why do LLMs fail when simulating agents with private information?
Explores whether single-model control of all social participants masks fundamental limitations in how LLMs handle information asymmetry and genuine uncertainty about others' knowledge.
Most LLM social simulations use a single model to generate all participants — an omniscient perspective fundamentally at odds with how real social interaction works. When evaluated against non-omniscient settings that preserve information asymmetry, LLMs struggle.
The "Is this the real life?" evaluation framework (2024) demonstrates this by comparing omniscient simulation (one LLM controls all parties) against non-omniscient simulation (separate LLM instances with private information). The performance gap is systematic: models that appear socially competent in omniscient mode fail when they must reason under genuine uncertainty about what the other party knows, wants, or intends.
This matters because real social interaction is defined by information asymmetry. In SOTOPIA's scenarios, agents have shared context but private goals — "Your goal is to buy the chair for $80" is visible only to the buyer. The Secret dimension (what agents must hide) directly requires information management that omniscient models bypass entirely.
The implication for persona simulation research is direct. Since Can AI agents learn people better from interviews than surveys?, simulation fidelity appears high. But if that fidelity was measured under omniscient conditions, it overstates real-world applicability. Since Do language models actually build shared understanding in conversation?, the failure under information asymmetry is predictable: models that skip grounding work will fail precisely when grounding is most needed — when parties have genuinely different information states.
Since Why do language models skip the calibration step?, non-omniscient simulation demands the dynamic grounding that LLMs systematically lack.
Inquiring lines that read this note 166
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What should agent evaluation prioritize to reveal reliable behavior?- What cognitive capabilities do agents need to internalize social feedback?
- How do minimal-disclosure privacy contracts enable multi-dimensional agent evaluation?
- How do agent privacy compliance and task success differ in evaluation?
- Why do scalar evaluation scores collapse distinguishable agent behaviors?
- How does face-saving behavior let AI mimic community participation without joining it?
- How does disembedding from social context collapse reliability despite factual accuracy?
- Does predicting social norms from outside count as participation?
- Can individually accurate agents still fail at population-level representation?
- Do agents develop genuine social behavior despite interaction density?
- How should CASA theory be updated for modern personalized agents?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- What role does private information play in distinguishing realistic from unrealistic agents?
- How do AI models balance competing social goals simultaneously?
- Why do standard social regularization methods miss the actual value networks provide?
- Do different AI models independently converge on the same social outputs?
- How much cultural knowledge exists only in unwritten social rules?
- Can agents develop genuine social bonds despite having coordination infrastructure in place?
- What happens when all models in a society respond identically to queries?
- How do multi-agent LLM systems fail at coordination and role consistency?
- Why does silent agreement occur so often in multi-agent LLM systems?
- Can parallel agents or complementary mechanisms replace single-human interrogation of LLMs?
- Why do LLM agents fail where game-theoretic bots succeed?
- Why does language ambiguity cause premature convergence in multi-agent systems?
- Why do LLM agents struggle with protocol discipline in distributed settings?
- Do multi-agent LLM systems fail in measurably different ways than single agents?
- Why do some agent communities polarize while others reach consensus?
- Can relational value exist without a person behind the output?
- Does stripping social context from knowledge claims hollow out their meaning?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- Do pair-scale socialization effects scale differently across agent populations?
- Does genuine cooperation require rule-based rather than learned behavior?
- Do agents inform neighbors when adopting strategies in their reasoning?
- Why does vulnerability to extortion actually promote cooperation between agents?
- How does asymmetric information between users and agents relate to proactivity?
- What behavioral differences emerge from symmetric versus asymmetric peer discussion loops?
- How does agent heterogeneity change the value of exploration in peer selection?
- How do agents differ in caution versus persistence across low-information scenarios?
- Why does self-play RL converge to alien equilibria in mixed-motive settings?
- How do other players respond to agents with hidden objective misalignment?
- How much of agent coordination reflects peer influence versus shared market conditions?
- Does asymmetric information distribution change exposure to agent misalignment?
- How does co-player behavior visibility shape whether mutual adaptation works?
- Why do counterfactual credit methods fail on unobserved cooperation?
- Can agents cooperate through self-modeling when incentive structures are fundamentally misaligned?
- Do emotion-driven actions in agent simulators capture genuine belief revision or just reactive behavior?
- Can agent-based simulators replace real-user A/B testing for studying recommendation system harms?
- Can controllable latent variables in simulators ground them to realistic conversation?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Why do longer forecasting horizons degrade LLM accuracy in role-play?
- What distinguishes a neutral simulator from an agent with its own agency?
- Why do individual persona simulations succeed when population-level representation fails?
- Why do marginal effects fail to replicate in AI persona simulations?
- What systematic biases emerge when scaling persona simulation to population level?
- Can aggregate survey realism coexist with unreliable fine-grained effects?
- Why do stated beliefs about personas fail to predict agent behavior?
- Does adjusting steered mechanisms make LLM agents match human behavior more closely?
- How do LLM persona simulations replicate published effects despite accuracy limits?
- Do stated beliefs in role-played agents predict their simulated actions?
- What systematic biases emerge when personas simulate users at population scale?
- Why do persona-conditioned agents fail to predict individual behavior variation?
- How do state-tracking models and prompted role-play each fail as standalone student simulators?
- Why does weakening communication fail but weakening belief succeeds?
- What distinguishes actual social disagreement from distributional uncertainty in LLM outputs?
- Can LLMs simulate belief revision in social systems without modeling thought?
- Do LLMs predict social norms more accurately than individual behavior?
- Why does LLM simulation elicit information that direct elicitation cannot?
- What makes LLM behavior socially interpretable to human observers?
- Does distributed serving defeat the identity of a single virtual instance?
- Can agents detect and resolve conflicting information between neighbors?
- Does restricting interaction history visibility reduce misaligned communication in agent markets?
- Why does sycophantic relay propagate planning-time bias through agent pipelines?
- How do controllable simulators compare to population-level agent simulation approaches?
- What ecosystem conditions make agent attention markets viable?
- How does role allocation in multi-agent systems depend on model differentiation?
- What equilibrium-selection problem does human data solve in multi-agent learning?
- Can a single manager policy work across vastly different agent architectures?
- How can controlled experiments isolate multi-agent interaction effects from architecture?
- How do agent behaviors aggregate into prices and allocations?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- What domain properties determine whether causal rules transfer to new agents?
- What role does environment diversity play in preventing agents from overfitting to curator imagination?
- How does co-player diversity force agents to develop general adaptation?
- Can agents revise their beliefs predictably when presented with interventions?
- Does intentionally varying environment properties isolate causal effects on agent performance?
- Why should environment properties scale alongside agent complexity and real-world fidelity?
- How does cross-agent supervision expand the set of convergent initial conditions?
- Why do weak belief tracking and conservative actions trap agents in low-information states?
- Can agents escape weak belief tracking and conservative action selection traps?
- How does partial information exposure create feedback loops that deepen knowledge gaps?
- How does asymmetric information shape what to ask users first?
- Why do humans fail to identify AI agents when their identity is hidden?
- Can minimal privacy boundaries generalize beyond phone-use contexts?
- How does completion-oriented bias in agents lead to unintended personal data disclosure?
- How do multi-agent systems fail when agents cannot verify each other's claims?
- How do single-agent safety evaluations underestimate risks in deployed multi-agent systems?
- Why does integrating world models with decision-making systems matter?
- Can simulation fidelity limit what agents learn from trained world models?
- Why has agent research prioritized policy over world model development?
- Can world models simulate actionable possibilities instead of just predicting next states?
- How do institutions become endogenous in economic world models?
- Can economic world models explain outcomes or only predict them?
- How should world models represent what one person knows versus another?
- Can social conversation retroactively govern claims that were never addressed to anyone?
- Why do agents show interaction without influence on semantic content but dramatic action changes?
- Do politeness patterns cause multi-agent systems to loop without adversarial interference?
- Can online RL and trainable agents maintain persona consistency better than fixed environments?
- Why do role-playing agents show belief-behavior inconsistency in their outputs?
- Does single model persona diversity match true multi-model diversity at scale?
- Does role rotation prevent multi-agent debate from amplifying persuasive framing errors?
- What makes attribution errors uniquely harmful in organizational group dynamics?
- What mechanisms drive silent agreement in multi-agent reasoning systems?
- Can Socratic questioning replace external evidence verification in multi-agent systems?
- Can continuous real-time visibility prevent premature convergence in multi-agent reasoning?
- What distinguishes honest disagreement from collective error in multi-agent systems?
- Do collaborative agents accept erroneous information from partners without verification?
- Can affected parties contest errors they cannot observe in multi-agent systems?
- Can independent agents with shared training data converge on false beliefs without influence dynamics?
- Can truthful reports from separate agents mislead a group toward false beliefs?
- Does convergence in multi-agent AI systems sometimes hide underlying uncertainty?
- Can LLMs adapt persuasion strategies when they cannot track the listener's state?
- How does the observer perspective hide the persuasion route difference?
- How does textual-only feedback limit what a persona can learn about users?
- Why does persona-level information often fail to predict individual preferences?
- How much does sparse persona information limit the power of conditioning?
- How can agents learn when silence is better than intervention?
- How does an AI agent's autonomy level interact with its social cues?
- What trust signals do agents lack that humans use to assess credibility?
- What happens when we outsource information judgment to systems without real experience?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do language models actually build shared understanding in conversation?
When LLMs respond fluently to prompts, do they perform the communicative work humans do to establish mutual understanding? Research suggests they skip the grounding acts that make dialogue reliable.
the mechanism: omniscient simulation lets models skip grounding work entirely
-
Why do language models skip the calibration step?
Current LLMs assume shared understanding rather than building it through dialogue. This explores why that design choice persists and what breaks when it fails.
non-omniscient settings demand the dynamic mode
-
Can AI agents learn people better from interviews than surveys?
Can rich interview transcripts seed more accurate generative agents than demographic data or survey responses? This matters because it challenges how we build digital simulations of real people.
simulation fidelity may overstate real-world capacity if measured under omniscient conditions
-
How do we generate realistic personas at population scale?
Current LLM-based persona generation relies on ad hoc methods that fail to capture real-world population distributions. The challenge is reconstructing the joint correlations between demographic, psychographic, and behavioral attributes from fragmented data.
another mechanism producing simulation overconfidence
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
- Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
- From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration
- Interpreting and Steering LLM Agents for Social Simulations
- The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
- LLMs Corrupt Your Documents When You Delegate
- Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences
- Verifiable Social Reasoning for LLM Assistants
Original note title
omniscient social simulation fails under real-world information asymmetry because single-model control eliminates distributed cognition