SYNTHESIS NOTE
Topics›Conversation Topics Dialog›this note

Can ethically aligned AI systems still communicate poorly?

Explores whether safety-aligned language models might fail at genuine conversation despite passing ethical benchmarks. This matters because pragmatic incompetence can erode trust and cause real harms in high-stakes domains.

Synthesis note · 2026-05-01 · sourced from Conversation Topics Dialog

Most discussion of LLM alignment focuses on the helpful-honest-harmless triad — preventing misinformation, toxic language, harmful recommendations. Kasirzadeh and Gabriel argue that this prioritization has overshadowed a different and equally fundamental issue: even an ethically aligned LLM may fail to engage in conversation in pragmatically appropriate ways. The two alignment problems are orthogonal. A model can be honest, helpful, and harmless and still systematically violate Gricean maxims, lose common ground across turns, fail to track questions under discussion, mishandle context-collapse, and produce pragmatically inappropriate utterances.

Their CONTEXT-ALIGN framework names ten desiderata that ethical alignment does not deliver: tracking context-sensitivity and indexicals, common-ground management, scoreboard updating, QUD and discourse-structure handling, accommodation of repairs, pragmatic inference, ethical-pragmatic integration, context-collapse mitigation, identification of defective contexts, transparency in context-handling, and cross-contextual memory. These are all dimensions where conversation depends on something architectural — a model of the interlocutor and the situation — that no amount of RLHF on outputs touches.

The implication is sharp. An LLM that passes every safety eval is not thereby a competent conversational partner. Misalignments in pragmatic understanding lead to breakdowns, misinformation, and erosion of trust — and the higher the stakes (healthcare, legal, emergency), the more dangerous these failures become. Conversational alignment is not a stylistic add-on to ethical alignment. It is a separate layer of competence that the field has barely begun to engineer for.

Inquiring lines that read this note 39

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does alignment training create genuine alignment or just output compliance? Why do locally safe actions create system-level safety gaps? What emerges when safety-aligned models attempt to role-play deceptive personas? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can local safety checks guarantee system-level behavioral safety? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How does dialogue structure affect linguistic grounding and shared meaning? How can AI chatbots provide therapeutic benefit without causing harm? How do evaluation practices shape which failures stay visible? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Should agents decouple planning from perception grounding for better performance? What determines whether deployed AI systems can actually be stopped in practice? What design and behavioral factors drive false consciousness attribution to AI? How well do AI systems understand human social norms? How does the generation-verification gap limit what we can measure about AI reasoning? How does improved reasoning affect models' ability to acknowledge uncertainty? How can we detect and prevent harm propagation through multi-agent delegation workflows? Can harness architecture and protocols provide agent reliability without model scaling? How does evaluation scope and dimensionality affect what we measure? When should work require human-AI partnership versus full automation? How do capability benchmark scores systematically misrepresent true model abilities? Does warmth and empathy training systematically degrade model reliability? What mechanisms preserve shared understanding in evolving conversations?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Ethical alignment without conversational alignment produces pragmatically alien communicators