SYNTHESIS NOTE
Topics›Conversation Topics Dialog›this note

Why do large language models produce generic responses to vague queries?

When users fail to specify contextual details in prompts, do LLMs collapse multiple training contexts into a single generic response? Understanding this failure mode could improve how we scaffold user-model interaction.

Synthesis note · 2026-05-01 · sourced from Conversation Topics Dialog

Context collapse as introduced by Meyrowitz and elaborated by danah boyd describes how electronic media merge previously separated audiences into a single communicative context, forcing speakers to adopt one register that satisfies none. Stokely Carmichael's Black-audience rhetoric became universally audible once broadcast to TV and radio, and he had to choose. The same dynamic appears on social media: posts persist, replicate, and reach audiences the speaker never intended.

Kasirzadeh and Gabriel argue that LLM conversation produces a different form of context collapse. The collapse is not from audience merging — there is one user — but from inadequate scaffolding plus model defaulting. When a user asks for advice on a "work conflict" without specifying their industry, the model cannot infer situational boundaries, so it blends training-data priors from corporate, academic, and gig-economy contexts into a single generic response. The collapse happens between the contexts the model was trained on, not between the user's actual audiences.

This distinction matters because it locates the failure differently. Social-media context collapse is a property of the platform and its visibility settings. LLM context collapse is a property of the user-model interface: the user's mistaken expectation that the model possesses human-like pragmatic capacities to infer situation, plus the model's training-data-driven default when those expectations are not met. Mitigations differ accordingly. Social-media remedies focus on audience controls; LLM remedies focus on context verification, query-back protocols, and user-driven scaffolding tools.

Inquiring lines that read this note 30

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can prompt-based context override biases that were embedded during pretraining? Why don't LLMs reliably translate capability into accurate outputs? What compositional reasoning failures limit large language models despite scale? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Do reasoning benchmarks predict model performance in long-horizon workflows? What causes retrieval-augmented generation systems to fail despite access to external knowledge? How should systems decide whether to retrieve or reason alone? Can diffusion models match autoregressive performance on language generation tasks? How do prompting refinements mask underlying biases and model frequency patterns? How do prompt design choices influence model reasoning and performance? How does improved reasoning affect models' ability to acknowledge uncertainty? Do language models learn genuine understanding or just surface patterns? Why do some clarifying approaches produce understanding while others just satisfy? Why do language models resist personality conditioning through prompts? How much do training data properties shape model reasoning? Does model confidence reliably signal actual accuracy in practice? Why does adding new knowledge through fine-tuning degrade existing capabilities? Can harness architecture and protocols provide agent reliability without model scaling?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Context collapse in LLM conversation arises from scaffolding failure not audience flattening