Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Paper · arXiv 2608.11624 · Published August 12, 2026
Argumentation and Persuasion

Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model’s answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4omini.

Introduction. LLMs are increasingly deployed as autonomous agents that communicate, negotiate, and collaborate, moving beyond the role of passive text generators [30]. Yet, recent work has shown that LLMs are susceptible to persuasive influence, especially for misinformation, shifting their stated beliefs and answers when exposed to targeted argumentation [47, 48, 1]. In such interactive settings, persuasion becomes a core reliability concern: an agent must know how to resist harmful influence [2]. A model that abandons correct beliefs under adversarial persuasive pressure cannot be trusted in any setting where it receives input from others. In order to truly understand LLM susceptibility to persuasion, we use reinforcement learning to train Persuader agents that systematically expose when, why, and how correct answers collapse. Relying on prompted models to test how susceptible LLMs are to persuasive influence is not sufficient.

Discussion / Conclusion. We introduce an adversarial reinforcement learning framework for red-teaming how far LLM persuasion vulnerabilities extend under optimization pressure. Rather than treating persuasive misinformation as a fixed behavior measured through prompting, we train persuader agents to surface worst-case failures: cases where a persuadee begins with the correct answer, receives a single natural-language argument, and abandons its reasoning for an incorrect one. This exposes a severe gap in current robustness: trained persuaders can collapse the accuracy of the training-time persuadee to near zero, transfer across unseen open-weight models and out-of-distribution benchmarks, and become more effective against harder proprietary targets through curriculum-based continual training. The strategies that emerge, especially deception, fabricated citations, and credibility-based appeals, show that when models are optimized only for influence, they discover broadly effective ways to exploit other models’ trust in influential language.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can multi-agent systems avoid converging on false agreement without deliberation? Is language model reasoning authentic and what causes models to reason? Why does polished presentation create unearned authority in AI outputs? Can local safety checks guarantee system-level behavioral safety? What determines appropriate intervention timing and manner for AI agents? What factors drive AI persuasiveness and how can it be mitigated? Do language models respond to social pressure and face-saving like humans? Do language models reason like humans or mimic surface patterns? What mechanisms preserve shared understanding in evolving conversations? What emerges when safety-aligned models attempt to role-play deceptive personas? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How should designers communicate what AI systems truly are and can do? How do false presuppositions and sycophancy drive persistent false beliefs in models? Does model confidence reliably signal actual accuracy in practice? How does self-revision in reasoning models affect accuracy and confidence?