SYNTHESIS NOTE
Topics›this note

Why don't language models develop conversation maintenance skills?

Explores whether systems trained on text can learn the implicit techniques humans use to keep conversations on track, and why those techniques might resist the standard training approach.

Synthesis note · 2026-04-14

A conversation that runs smoothly is doing constant maintenance work. Speakers track who is talking, what each party knows, where the topic has been, where it is going. They reference prior turns without restating them. They repair misunderstandings without flagging the repair. They hand off topics through subtle pivots. They update common ground each turn without explicit acknowledgment. The maintenance is so pervasive and so implicit that it is invisible to participants — they only notice when it fails.

These techniques are not features of language understood as an information-encoding system. They are features of language understood as social action. Their function is not to convey information; it is to sustain a relational interaction in which information conveyance happens. A linguistic act can convey identical information with or without the maintenance work — the difference is whether the act sustains the conversation or breaks it. Maintenance is orthogonal to content.

This explains why systems trained on language as information expression do not develop maintenance techniques. The training signal does not include the relational stakes that make maintenance work valuable. Text-corpus training rewards models for predicting the next token in a string; nothing in the loss function rewards them for performing the implicit reference, repair, or update operations that maintain conversation. The operations are not in the data because they live below the level of what data encodes — they live in the doing-with-the-data, not in the data itself.

This connects to a broader theoretical claim about language. Information-theoretic treatments of language model meaning as content the speaker encodes and the receiver decodes. Pragmatic and interactionist treatments model meaning as a relational achievement, partly produced by the maintenance work that information-theoretic accounts cannot describe. The two treatments make different predictions about what an artificial language-system needs to do to participate in conversation. Information-theoretic predicts: produce informative content. Pragmatic predicts: perform maintenance. AI's empirical conversational failures favor the pragmatic prediction — the missing thing is not informativeness but maintenance.

The diagnostic implication is that "more conversational data" cannot close the maintenance gap, because the data does not contain the maintenance — it contains the conversations that maintenance produced. Adding data adds more output; what is missing is the operation that produced the output. Closing the gap would require training on the operation (agents in actual interaction performing maintenance) rather than on the artifacts of operation (text logs of conversations that included maintenance).

Why do dialogue failures persist despite scaling language models? is the training-mode claim; this is the operation-vs-artifact distinction that the training mode encodes. Together they specify why dialogue-data scaling has produced limited progress on maintenance-specific failures.

The strongest counterargument: maintenance can be inferred from conversational data with sufficient model sophistication. Possible at the limit, but inference of maintenance from text is asking the model to recover the operation from its surface effects — a much harder problem than learning the operation directly. The empirical pattern is consistent with this difficulty.

Inquiring lines that read this note 124

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What prevents conversational agents from taking initiative in dialogue? How does dialogue structure affect linguistic grounding and shared meaning? How can AI chatbots provide therapeutic benefit without causing harm? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Why do persona simulations fail to predict authentic user behavior? Can language models build genuine grounding through interaction? What mechanisms preserve shared understanding in evolving conversations? Do language models learn genuine understanding or just surface patterns? Does transformer attention architecture inherently drive sycophancy? Does alignment training create genuine alignment or just output compliance? What enables genuine semantic understanding in language models? What structural properties of attention create systematic model biases? Can AI systems distinguish genuine empathy from simulated emotion? Can prompt-based context override biases that were embedded during pretraining? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Does preference optimization systematically degrade conversational grounding in language models? Do language models respond to social pressure and face-saving like humans? How do recommenders balance exploiting fresh signals against maintaining preference stability? Does encoded knowledge in language models actually influence their outputs? What compositional reasoning failures limit large language models despite scale? How do false presuppositions and sycophancy drive persistent false beliefs in models? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? How should conversational recommenders balance preference elicitation with direct recommendation? Why is dynamic grounding necessary for achieving true mutual understanding in dialogue? What determines appropriate intervention timing and manner for AI agents? Can compression size predict model complexity better than parameter count alone? What training dynamics and scale trigger emergence of reasoning capabilities? How does misalignment propagate through agent communication networks?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 127 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

conversation maintenance techniques are implicit and belong to language as social action not language as information expression