CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human–human interaction. However, human–AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human–chatbot dialogue across a range of chatbot behaviors and use cases. We release COMPAN- IONSIM: a simulation framework with 2,240 simulated human–chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample (N1 = 628) and Study 2 across the U.S., U.K., India, and Nigeria (N2 = 3, 646). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy.
Introduction. Hundreds of millions of people now use artificial intelligence chatbots based on large language models (LLMs) in functional contexts, including writing (Hill 2025), programming (Staff 2025), crafting emails (Buchanan and Paris 2025), and seeking medical advice (Rosenbluth and Astor 2025). Many chatbots are designed and marketed as “companions” or “friends” (Luka 2025), and general-purpose chatbots, such as ChatGPT, have come to be viewed as companions by many users (Manoli et al. 2026). These relationships center on the humanlike characteristics of chatbots that evoke anthropomorphism, such as expressions of emotion that have been shown to increase social presence (Konya- Baumbach, Biller, and Von Janda 2023) and facilitate social bonds (Maeda and Quan-Haase 2024). Prior work has argued that humanlike design facilitates user engagement with technological systems (Chan- dra, Shirish, and Srivastava 2022); anthropomorphic systems can provide benefits to users, such as suggesting patterns of socially appropriate interaction (Złotowski et al.
Discussion / Conclusion. We developed COMPANIONSIM to support the rapidly growing interest in studying human–AI companionship interactions. Across two proof-of-concept empirical studies, we find that companionship behaviors play a significant role in third-party chatbot perceptions, and we provide evidence for the methodological viability of dynamic LLM simulations. Our results suggest that behavior-controlled simulations and synthetic data could help with the challenges of developing benchmarks and other tools for AI safety. 5.1 Need for Disaggregated Measurement AI use frequency, age, and gender predicted the effects of companionship behavior on likability, humanlikeness, affective trust, and cognitive trust. These findings extend findings on the overall effects of humanlike model behavior on user perceptions or variation across a single demographic variable (e.g., gender) (Ibrahim et al. 2025; Sharma et al. 2023; Chaves et al. 2022).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What drives appropriate trust calibration in personalized AI systems?- Why does conversational style make ChatGPT seem more trustworthy to users?
- How does outcome feedback change beliefs about AI versus human partner reliability?
- How does understanding persistent journeys intensify both trust and privacy concerns?
- Which specific chatbot behaviors drove the drop in likability and trust ratings?
- How do chatbot design features like intimacy-by-design sustain romantic bonds?
- What distinguishes romantic chatbot bonds from other forms of AI companionship?
- Do users consciously recognize their needs before forming chatbot relationships?
- Do AI companions reduce loneliness compared to talking with another person?
- Can perceived understanding from a chatbot exist alongside feeling alone?
- How does emotional dependence on chatbots affect user wellbeing?
- How do dropout rates and low adherence affect chatbot therapy outcomes?
- How do user expectations change as chatbots remember more interactions?
- How does the expectation ratchet affect long-term chatbot satisfaction?
- What temporal design dimensions characterize different chatbot relationship types?
- How do time gaps between conversations change what chatbots should remember?
- Does personalization help or hurt persistent companion chatbots?