SYNTHESIS NOTE
Topics›Theory of Mind›this note

Do LLMs predict persuasion based on actual dialogue or training bias?

Why do large language models consistently predict concession-based persuasion intentions even when dialogue context suggests otherwise? Understanding this gap reveals how alignment training shapes not just model behavior but also how models perceive others' intentions.

Synthesis note · 2026-02-22 · sourced from Theory of Mind

When asked to infer persuasion intentions from dialogue, most LLMs exhibit a systematic bias: they predict intentions "characterized by making the other person feel accepted through concessions, promises, or benefits" — regardless of whether the actual dialogue context supports this inference.

The hypothesis is that RLHF (Reinforcement Learning from Human Feedback) is the mechanism. RLHF "tends to prioritize safety and politeness" during preference optimization, and this training signal bleeds into intention prediction. The model has learned that conciliatory, benefit-oriented responses are preferred by human raters, and this preference leaks into its predictions about what other agents will do — it projects its own trained disposition onto the agents it's modeling.

This is a specific, measurable instance of a broader pattern: alignment training shapes not just what the model says but how it models others. If RLHF teaches the model that accommodation is preferred, the model begins to assume accommodation is what agents do. It becomes harder for the model to represent genuinely adversarial, manipulative, or hardball persuasion strategies because its own training bias makes these strategies less probable in its prediction space.

The practical consequence for persuasion-aware AI: a model biased toward predicting concessions will systematically underestimate adversarial intent. In negotiation support, threat detection, or social manipulation detection, this bias translates directly into blind spots — the model expects cooperation where exploitation is occurring.

Inquiring lines that read this note 66

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What factors drive AI persuasiveness and how can it be mitigated? How do prompt design choices influence model reasoning and performance? Do language models reason like humans or mimic surface patterns? Is language model reasoning authentic and what causes models to reason? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can mechanistic interpretability reliably guide practical model design choices? Do language models respond to social pressure and face-saving like humans? How do false presuppositions and sycophancy drive persistent false beliefs in models? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? What prevents conversational agents from taking initiative in dialogue? What emerges when safety-aligned models attempt to role-play deceptive personas? How does dialogue structure affect linguistic grounding and shared meaning? Can prompt-based context override biases that were embedded during pretraining? What determines whether deployed AI systems can actually be stopped in practice? Can language models build genuine grounding through interaction? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What makes personas effective for predicting individual preferences and behavior? How do recommenders balance exploiting fresh signals against maintaining preference stability? Does alignment training create genuine alignment or just output compliance? How do training data properties determine the emergence of internal misalignment? What linguistic features distinguish AI-generated text from human writing most reliably? What types of diversity prevent reasoning systems from collapsing? Why do persona simulations fail to predict authentic user behavior? Do language models learn genuine understanding or just surface patterns? Does transformer attention architecture inherently drive sycophancy?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 184 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

RLHF biases LLMs toward predicting concession-based persuasion intentions regardless of dialogue context