Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Paper · arXiv 2607.27816 · Published July 30, 2026
Role-Play and Persona Behavior

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA’s output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles.

Introduction. Role-playing agents (RPAs) provide interactive storytelling, companionship, and emotionally engaging conversation. Their value often unfolds over multiple turns, and deployed systems already serve millions of users (Irvine et al., 2023). As these systems enter large-scale real-world use, reliable evaluation becomes essential for measuring capability, comparing systems, and guiding further improvement. Unlike general-purpose task assistants, the object being evaluated is not an isolated response but a dialogue jointly constructed by the RPA and its user. Such dialogues rarely have a uniquely correct answer, and conventional automatic metrics correlate poorly with human judgments in open-ended dialogue (Liu et al., 2016). Quality must instead be established through measures of character fidelity, narrative development, and user experience.

Discussion / Conclusion. Fixed-history evaluation places an RPA in a trajectory that it did not help construct, thereby mixing its own capability with the influence of the external history. User-independent scoring further compresses genuine individual differences in satisfaction into a single standard. Palate uses simulators trained from real user histories to generate free multi-turn dialogues and evaluates candidates along personalized, generic turn-level, and whole-session tracks. It further decomposes one person into a behavior policy learned from real dialogue and an individual utility supervised by experience annotations, changing the basic unit of role-play evaluation from an isolated RPA to a user–RPA pair. This introduces a user-perspective satisfaction reference while retaining general quality evaluation. Across 16 candidates, advantages on the three tracks do not coincide, and the five users do not share a single best candidate. The central output of Palate is therefore not another scalar leaderboard, but an interactive evaluation profile that locates cross-track capability mismatches and per-user differences.

Lines of inquiry this paper opens 14

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do agent-learned skills transfer and improve across different tasks? Why do persona simulations fail to predict authentic user behavior? How can reward models capture diverse human preferences without excluding minority populations? Why do standard benchmarks fail to predict agent deployment success? How do capability benchmark scores systematically misrepresent true model abilities? Do reasoning benchmarks predict model performance in long-horizon workflows? What design and behavioral factors drive false consciousness attribution to AI? How do we enforce security boundaries in evaluation environments? Do language models reason like humans or mimic surface patterns? How well do AI systems understand human social norms? What emerges when safety-aligned models attempt to role-play deceptive personas?