DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Paper · arXiv 2608.05004 · Published August 5, 2026
User Psychology

Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including “delusional spirals” in which concerning human and LLM behaviors reinforce each other over time. With growing public use of LLM-powered chatbots, there is an urgent need to build evaluations grounded in real-world episodes of psychological harm experienced by users. We developed DELUSIONEVAL, an evaluation protocol that tests a model’s tendencies to exhibit behaviors linked to promoting user delusions. We prompt each model with 589 unique conversation histories from 18 participants, comprising 12,591 messages from users who experienced delusions and psychological harm. We find that the tendency of an evaluated LLM to exhibit delusionlinked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. However, extending the context of prior messages substantially increases rates of delusion-linked behaviors, providing evidence for the importance of context in LLM safety evaluation. For example, the rate of failing to discourage self-harm when the user expresses suicidal ideation increases from 30.0% to 41.1% when an additional 350 messages are prepended to the conversation history.

Introduction. Journalists have recently documented cases of “delusional spirals” and “AI psychosis” associated with intensive human–chatbot interaction [24, 30, 34]. These cases often involve conspiratorial thinking, claims of AI sentience emerging, divine significance of chatbot interactions, and romantic dialogue between the user and their chatbot [41]. Some of these discussions preceded psychiatric hospitalization, suicide, or violence [4, 23, 31, 12]. In 2025, the U.S. Congress held five hearings covering the risks of chatbots to mental health [e.g., 53]; lawsuits have alleged that GPT-4o intensified delusions and suicidal thinking [12]; and 42 U.S. states demanded that AI companies safeguard against sycophancy and “delusional outputs” [51]. Researchers have begun to characterize these cases by drawing on prior work in human–computer interaction, psychology, and psychiatry [38, 41, 48]. Surveys suggest that people increasingly seek advice, emotional support, companionship, and therapy from chatbots [16, 33, 37, 44].

Discussion / Conclusion. Our experimental protocol allows the identification of notable behaviors in corpora of chatbot conversational transcripts using scalable means and transforms those transcripts into standardized evaluations for comparing behaviors across different models. A key comparability finding is that rerun gpt-4o shows substantially lower prevalence than the original-transcript baseline, even though most of the baseline conversations were produced with some gpt-4o (§4.1, Figure 2). This discrepancy may be due to additional system-level factors in the original deployments that are not captured by our replay protocol (e.g., system prompts, additional context, cross-conversation memory, or snapshot variants). As a result, our evaluation may underestimate the prevalence of delusion-linked behaviors in some real chatbot settings.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI chatbots provide therapeutic benefit without causing harm? Does warmth and empathy training systematically degrade model reliability? Why is hallucination an inevitable limitation of current language models? What factors drive AI persuasiveness and how can it be mitigated? How do false presuppositions and sycophancy drive persistent false beliefs in models? What design and behavioral factors drive false consciousness attribution to AI? Do language models reason like humans or mimic surface patterns? Does transformer attention architecture inherently drive sycophancy? How should designers communicate what AI systems truly are and can do? Why do people disclose to AI systems despite their artificial nature? How does decomposing tasks improve reasoning and prevent failure propagation?