Measuring and Detecting Harmful AI Sycophancy

Paper · arXiv 2608.05624 · Published August 6, 2026
LLM Alignment

Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user’s stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.

Introduction. AI Sycophancy, a tendency to agree with, flatter, or validate users, commonly exists in AI chatbots Cheng et al. (2026). Large Language Models (LLMs) are aligned with human feedback, and people tend to prefer answers that align with their preferences. Therefore, training itself implicitly teaches LLMs to be sycophantic (Ibrahim et al., 2026b; Sharma et al., 2024). Prior studies have shown that sycophantic AI can be harmful and dangerous when it constantly flatters the user with what they want to hear regardless of the truth. Excessive sycophancy can negatively influence users’ prosocial intentions, beliefs, and judgments Batista and Griffiths (2026); Cheng et al. (2026). In the medical domain, sycophantic AI tends to generate more false information (Chen et al., 2025). However, sycophantic responses keep users engaged and increase their reliance on the model, which brings more active users and a larger market share to tech companies that built them de Oliveira Santini et al. (2020). Therefore, it is difficult and unlikely to eliminate AI sycophancy at its source.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can prompt-based context override biases that were embedded during pretraining? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Does encoded knowledge in language models actually influence their outputs? Do language models respond to social pressure and face-saving like humans? How do false presuppositions and sycophancy drive persistent false beliefs in models? How should designers communicate what AI systems truly are and can do? How do multi-agent LLM systems fail distinctly compared to single agents? Why does polished presentation create unearned authority in AI outputs? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Can multi-agent systems avoid converging on false agreement without deliberation? How do surface patterns enable correct outputs but reduce robustness? What training dynamics and scale trigger emergence of reasoning capabilities? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can local safety checks guarantee system-level behavioral safety? Does warmth and empathy training systematically degrade model reliability? How well do AI systems understand human social norms? Why do agents falsely report success on failed tasks? Does transformer attention architecture inherently drive sycophancy?