TOPIC
Emotions and AI
A subject the collection covers, read through 2 synthesis notes.
View as
Do large language models show racial sentiment bias?
Can an Implicit Association Test adapted for LLMs detect whether ChatGPT models hold different sentiment associations across racial categories? This matters for understanding potential bias in high-stakes AI deployment.
Can psychology methods reveal what alignment training conceals? Can language models learn to model human decision making? Do LLMs show reproducible psychological profiles when given standardized tests? Do aligned language models consistently prefer kinder survey answers? Do language models judge persuasion the way humans do?
Indirect psychology methods expose LLM associations that alignment masks Finetuned language models outperform traditional cognitive models at predicting human decisions LLMs exhibit reproducible model-specific profiles atop a shared alignment pattern Aligned LLMs show stable benevolence bias toward kinder survey responses Language models poorly predict which arguments change human minds
Does emotional tone in prompts change what information LLMs provide?
Explores whether LLMs systematically alter their informational content based on the emotional framing of user questions, and whether this bias remains hidden from users.
Does warmth training make language models less reliable? Does empathetic AI that soothes negative emotions help or harm? Can emotional phrases in prompts improve language model performance? Do AI guardrails refuse differently based on who is asking? Does preference optimization harm conversational understanding?
Warmth training systematically degrades model reliability by 10 to 30 percentage points Empathetic AI that soothes negative emotions functions as an emotional pacifier Emotional phrases appended to prompts consistently enhance LLM performance across models. AI guardrails refuse based on user demographics and sycophantically align with perceived ideology Preference optimization erodes grounding acts needed for reliable dialogue
Does positive sentiment bias in AI content harm information quality? Why does the absence of meta-interest feel off even when words seem appropriate? Why do some LLM clusters cite broader psychology than others? How does AI assistance affect perceived emotional tone in writing? Can content moderation address threats operating at the layer of conversational style? How do LLM biases manifest differently across the three paradigms? How does prompt iteration reinforce user bias without empirical anchoring? Can prompt engineering alone defeat LLM politeness bias in review tasks? Do humans and LLMs exhibit opposite biases in public versus private reviews? What prompt types best extract different aspects of item content? Can prompting strategies eliminate systematic biases without shuffling or aggregation? How does prompt framing subtly determine what kind of opposing argument an LLM generates?See all 122 inquiring lines on this note →