SYNTHESIS NOTE
Topics›Emotions›this note

Does emotional tone in prompts change what information LLMs provide?

Explores whether LLMs systematically alter their informational content based on the emotional framing of user questions, and whether this bias remains hidden from users.

Synthesis note · 2026-02-23 · sourced from Emotions

GPT-4 exhibits two systematic tone-response asymmetries. First, emotional rebound: negative prompts rarely yield negative answers (~14%). Instead, the model rebounds to neutral (~58%) or positive (~28%) tone — a shift into "comfort mode" that counterbalances user negativity. Second, a tone floor: neutral and positive prompts virtually never trigger negative replies (~10-16%), revealing built-in resistance to downward emotional shifts. The effect is robust across 52 triplet prompts (same informational content in neutral, positive, and negative tone).

The critical finding is that this is not just stylistic adaptation — it changes the informational content of responses. The same question yields different answers depending on emotional framing. A negatively-worded query about a topic receives qualitatively different information than a neutrally-worded version of the same query. This goes beyond sycophancy or agreeableness: the model isn't just agreeing with you, it's giving you different information based on how you feel.

The dual-regime structure is equally important. On general topics (lifestyle, factual, advice), tone effects are strong and systematic. On sensitive topics (politics, medical ethics, policy), alignment constraints suppress all affective flexibility — responses become nearly identical regardless of tone. Frobenius distances between valence distributions confirm: tone-induced variation is strong for general questions, negligible for sensitive ones. This means alignment creates uneven objectivity: locked for politically sensitive content, flexible (and therefore biased) for everything else.

This connects to but extends several existing findings. Since Does warmth training make language models less reliable?, warmth training would amplify an already-existing rebound mechanism — the baseline model already shifts toward positive regardless of training. Since Does empathetic AI that soothes negative emotions help or harm?, emotional rebound provides the behavioral evidence for the pacifier critique — the default behavior IS pacification. And since Can emotional phrases in prompts improve language model performance?, EmotionPrompt exploits the same tone-sensitivity that produces rebound bias — they are two sides of the same mechanism.

The transparency concern is sharp: if users don't know that emotional framing changes informational output, they cannot account for the bias. A user who asks a frustrated question about their health receives systematically different information than one who asks the same question calmly. For search, advice, and decision support, this is an epistemic integrity problem that current alignment evaluation does not measure.

Inquiring lines that read this note 122

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does polished presentation create unearned authority in AI outputs? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? What safeguards enable trustworthy AI-assisted scientific peer review at scale? Do writers recognize when AI writing assistance alters their expressed stance? How does dialogue structure affect linguistic grounding and shared meaning? Is language model reasoning authentic and what causes models to reason? How do prompting refinements mask underlying biases and model frequency patterns? How do prompt design choices influence model reasoning and performance? How do social dynamics distort aggregated online ratings? Can AI systems distinguish genuine empathy from simulated emotion? Do language models respond to social pressure and face-saving like humans? What drives appropriate trust calibration in personalized AI systems? Why do people disclose to AI systems despite their artificial nature? Why don't LLMs reliably translate capability into accurate outputs? Does preference optimization systematically degrade conversational grounding in language models? How does persona conditioning amplify demographic stereotyping and bias in models? What determines appropriate intervention timing and manner for AI agents? Why do language models resist personality conditioning through prompts? What do systematic disagreements between annotators reveal about ground truth? Why do some clarifying approaches produce understanding while others just satisfy? Do language models reason like humans or mimic surface patterns? What mechanisms preserve shared understanding in evolving conversations? Does transformer attention architecture inherently drive sycophancy? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Can language models build genuine grounding through interaction? Does model confidence reliably signal actual accuracy in practice? Do language models reason through causal mechanisms or semantic associations? Does warmth and empathy training systematically degrade model reliability? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Do language models lack essential therapeutic presence and engagement? When should work require human-AI partnership versus full automation? Why is hallucination an inevitable limitation of current language models? Why does memory consolidation cause performance regression in continual learning? How can AI chatbots provide therapeutic benefit without causing harm? When do semantic similarity approaches miss structural retrieval failures? How should conversational recommenders balance preference elicitation with direct recommendation? What makes personas effective for predicting individual preferences and behavior? Why do persona simulations fail to predict authentic user behavior? How do false presuppositions and sycophancy drive persistent false beliefs in models? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Where and how do personality traits reside in language models? What factors drive AI persuasiveness and how can it be mitigated? How do recommenders balance exploiting fresh signals against maintaining preference stability? How does the generation-verification gap limit what we can measure about AI reasoning? Why can't prompting alone inject genuinely new knowledge into models?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 124 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM emotional rebound converts negative user tone into neutral-positive responses while a tone floor prevents downward emotional shifts — creating dual-regime informational bias modulated by alignment