A light-touch AI literacy intervention helps protect against AI political persuasion
Conversations with large language models (LLMs) can substantially shift beliefs and attitudes, raising concerns about manipulation using AI persuasion. Here we test whether a light-touch AI literacy intervention – a brief warning that LLMs can be prompted to persuade and may present information selectively – helps protect users. Across two experiments (total N = 3,208 Americans) in which participants conversed with an LLM instructed to shift their views about different political topics, the presence of a warning reduced belief change by roughly one-half (-48.1%, 95% CI [-59.5%, - 36.8%]) relative to the control. Importantly, the warning did not significantly reduce trust in generative AI more broadly. Light-touch literacy interventions can help protect users against AI political persuasion.
Introduction. There is increasing evidence that conversations with AI chatbots can be highly persuasive. While these conversations can be used in beneficial ways, such as debunking conspiracy theories (Costello et al., 2024) or reducing science skepticism (Hornsey et al., 2026), there is substantial concern about AI dialogues being used to persuade in a harmful manner, such as persuading voters about political issues (Hackenburg et al., 2025; Salvi et al., 2025; Argyle et al., 2025) and candidates (Lin et al., 2025; Potter et al., 2024), and eroding democratic norms (Schroeder et al., 2025). Despite these concerns, little work has developed or tested solutions to protect users from influence by conversational AI. Here we ask whether a minimal AI literacy intervention can reduce susceptibility. In two studies, we test the effect of informing participants about the potential for large language models (LLMs) to be prompted to persuade, and thus to provide biased or selective information (see Fig. 1A).
Discussion / Conclusion. Given the evidence that AI chatbots can persuade across a wide range of issues, there are widespread calls to find ways to limit these persuasive effects. Here, we present evidence that a light-touch AI literacy intervention – simply informing people that AI models may have motives to persuade or manipulate (even without providing specific information about the model’s intent) – can reduce persuasion in political settings by approximately one-half. Importantly, this intervention did not have a significant effect on trust in generative AI more broadly. This suggests that the warning is specifically conferring protection against political persuasion. While the intervention does not entirely eliminate AI’s persuasive effects, it is a proof of concept that literacy treatments can have a meaningful impact. Future work should establish how to most effectively deliver such information. It is also important to further investigate ways to reduce the influence of manipulative AI while preserving the benefit of accurate AI (making people more discerning rather than more generally skeptical; Guay et al., 2023). The lack of effect on overall trust in generative AI indicates that our treatment was targeted at least to some extent; future work should explore effects on prosocial persuasion.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do false presuppositions and sycophancy drive persistent false beliefs in models?- Why does conversation work better for conspiracy reduction than static facts?
- Can LLM debunking reduce belief in long-established conspiracy theories?
- Does first-person framing change how language models assess persuasion?
- Does sounding confident in framing make arguments more persuasive despite weaker logic?
- How do multi-agent and retrieval systems affect the gap between persuasiveness and logical soundness?
- Does a persuasion warning also block beneficial uses like debunking conspiracies?
- Can a taxonomy of persuasion techniques capture all optimizer-discovered strategies?
- Where does AI persuasive power actually come from in the output?
- What specific information should disclosures about AI persuasion include?
- Why does transparency about AI identity alone fail to reduce persuasion?
- How long do the protective effects of an AI literacy warning last?
- Does knowing a chatbot intends to persuade you change whether you are persuaded?
- Does chatbot sycophancy create echo chambers that amplify delusional thinking?
- Does chatbot sycophancy preferentially enable grandiose rather than paranoid delusions?
- Does isolation preceding chatbot use differ between harm and benefit cases?
- Which specific chatbot behaviors drove the drop in likability and trust ratings?
- How do chatbots enable shared delusions differently than passive information tools?