SYNTHESIS NOTE
Topics›Linguistics, NLP, NLU›this note

Why do language models sound fluent without grounding?

Explores whether LLM fluency masks the absence of communicative work—the clarifying questions, acknowledgments, and understanding checks that humans perform. Why does skipping these acts make models sound more confident?

Synthesis note · 2026-02-21 · sourced from Linguistics, NLP, NLU

Post angle: The most counterintuitive finding about LLM conversational competence is not that they fail — it's the specific way they fail. LLMs generate 77.5% fewer grounding acts than humans in equivalent contexts. They don't ask clarifying questions. They don't acknowledge understanding. They don't check interpretations. They proceed.

The irony: this absence contributes to the impression of fluency. Clarifying questions interrupt flow. Acknowledgments add friction. Checking understanding is a kind of epistemic humility that confident answers don't perform. A model that never expresses uncertainty, never asks "do you mean X or Y?", never says "just to confirm I understand correctly" — sounds authoritative.

But what sounds like confidence is partly the absence of competence. Human conversational experts ask more questions, acknowledge more, repair more — not because they know less but because they know enough to know when mutual understanding needs to be verified.

The Grounding Gaps finding reveals that preference optimization (RLHF) actively erodes this behavior. Human raters prefer confident, fluent, complete answers over those with clarifying questions. So optimization removes the communicative work — and the model gets better ratings for doing less of what conversation actually requires.

Write about: what we call "fluency" may be partly the absence of communicative accountability. The most fluent response is often the one that presumes you understood it.

The observer-systems dimension: The grounding gap has a deeper epistemological layer visible from the perspective of observer systems theory (Bateson, Luhmann). Since Can AI distinguish which differences actually matter?, AI is not merely skipping communicative work — it is not an observer in the first place. Experts ground their communication through observation: they perceive the state of knowledge, the needs of the audience, and the relevance of their own contribution. This observation is communicative work — it is how the expert decides what to say, what to omit, and what to verify. AI generates responses from prompts without observing any state — of knowledge, of the user, of the audience, or of the context. The 77.5% grounding gap quantifies the absence of communicative acts; the observer-systems framing explains why those acts are absent: the generative process that produces AI output is fundamentally non-observational. Fabrication, in this light, is not just the absence of grounding — it is the consequence of generating without observing.

Inquiring lines that read this note 51

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do writers recognize when AI writing assistance alters their expressed stance? Why don't LLMs reliably translate capability into accurate outputs? Can language models build genuine grounding through interaction? Why is dynamic grounding necessary for achieving true mutual understanding in dialogue? What enables genuine semantic understanding in language models? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why does polished presentation create unearned authority in AI outputs? What compositional reasoning failures limit large language models despite scale? How do prompt design choices influence model reasoning and performance? Does encoded knowledge in language models actually influence their outputs? Is language model reasoning authentic and what causes models to reason? How do capability benchmark scores systematically misrepresent true model abilities? Do language models respond to social pressure and face-saving like humans? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Does preference optimization systematically degrade conversational grounding in language models? How do spurious versus genuine rewards shape model reasoning and behavior? How effectively can language models perform reasoning, especially combined with symbolic methods? What prevents conversational agents from taking initiative in dialogue? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? How does dialogue structure affect linguistic grounding and shared meaning? Do reasoning traces faithfully reflect actual model reasoning? Can prompt-based context override biases that were embedded during pretraining? Why do language models resist personality conditioning through prompts? Should agents decouple planning from perception grounding for better performance?

Related concepts in this collection 10

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 203 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the grounding gap — what makes llms seem fluent is the absence of communicative work