Assessing mentalization in humans and large language models
Mentalization - the ability to infer others’ beliefs and intentions to guide one’s own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks.
Introduction. Mentalization is a cognitive process associated with interpreting the behaviour of others as the result of latent mental states, beliefs and emotions [1–3]. In humans, mentalization shapes decisionmaking in social contexts by predicting the actions of others and adjusting one’s own behaviour accordingly [4–10]. Evidence also has suggested the presence of mentalization in a select range of non-human animal species, including dogs, corvids, chimpanzees and gorillas [11–20] reflecting an evolved capacity for social intelligence. Beyond humans and other animals, a topic of significant recent development concerns whether generative artificial intelligence (AI) systems, such as large language models (LLMs) demonstrate behaviour consistent with mentalizing. Researchers have turned to experimental methods applied within the fields of psychology, cognitive science and neuroscience to uncover features of behaviour and reasoning abilities in LLMs [21–30].
Discussion / Conclusion. A growing interest concerns whether artificial intelligence systems understand the thoughts and beliefs of others, and use this information to guide their own actions. Previous studies have assessed theory-of-mind in LLMs by measuring choice accuracy in response to story-based prompts, obscuring the latent mechanisms underlying observed behaviour. Here we employ two behavioural tasks previously validated in human participants, the inspection game and rock-paper-scissors (RPS). These tasks provide a normative assessment of machine intelligence [65], and allow the examination of behaviour at specific depths of mentalizing [69]. Across two strategic interaction games, humans demonstrated both recursive and adaptive mentalization, replicating previous results and demonstrating the robustness of the tasks used [80, 106, 107]. On the other hand, LLMs demonstrated varying mentalization strategies across the two strategic games, with significant differences observed across model providers. Furthermore, LLMs prompted using Social Chain-of-Thought (SCoT) demonstrated robust improvements in performance, reflecting more sophisticated reasoning.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do persona simulations fail to predict authentic user behavior?- Can fitted strategy models distinguish genuine mental-state representation from learned game policies?
- Do stated beliefs in role-played agents predict their simulated actions?
- What distribution patterns appear across different theory-of-mind datasets?
- How does theory of mind predict who benefits from AI collaboration?
- Why do reasoning models perform worse on theory of mind tasks?
- What makes social reasoning fundamentally different from mathematical reasoning?
- Why does increasing reasoning not improve AI social reasoning performance?
- Can multi-agent metacognitive decomposition achieve human-level theory of mind?
- Can reasoning scaffolds help with nuanced judgment tasks like empathy?
- Why might social reasoning work differently than formal logical reasoning?
- What makes social reasoning fundamentally different from formal logical reasoning?
- Can structured theory of mind benchmarks measure genuine mental state reasoning?
- Why do reasoning models regress on some theory of mind tasks?
- How does theory of mind predict success in human-AI partnerships?