When should AI agents ask users instead of just searching?
Explores whether tool-enabled LLMs should probe users for clarification when uncertain, rather than silently chaining tool calls that drift from intent. Examines conversation analysis patterns as a formal alternative.
Tool-enabled LLMs have a structural problem: when they can't immediately answer a query, they chain tool calls (search, calculation, code execution) and each intermediate step is conditioned on the output of the previous step. The result is progressive divergence from the user's original intent. The more tools the model uses, the further it drifts.
Conversation Analysis (Schegloff, 2007) offers a formal alternative from human talk-in-interaction. When human speakers can't immediately provide the expected response, they don't silently think harder — they insert a new pair of utterances to bridge the gap. These "insert-expansions" serve three functions: clarifying intent ("Do you mean the downtown location?"), scoping responses ("Are you looking for something under $50?"), and enhancing appeal ("I should mention it also comes in blue").
The key move is the "user-as-a-tool" paradigm: instead of the model consulting external tools and accumulating drift, it consults the user. The user provides necessary details and refines their request. This replicates exactly the structure of human insert-expansions — post-first inserts recover from misunderstandings, pre-second inserts gather information needed to choose the right response.
The empirical evidence from recommendation tasks shows benefits from this approach. But the deeper point is architectural: since Why can't conversational AI agents take the initiative?, the insert-expansion framework gives a principled answer to WHEN agents should break passivity — not by adding unsolicited content, but by asking structured questions when their internal processing would otherwise diverge.
This connects to the distinction between formal and functional linguistic competence: LLMs have formal competence (handling language in itself) but lack functional competence (doing things WITH language — reasoning, using world knowledge, establishing common ground). Insert-expansions are a functional linguistic capability. The paper argues that natural speech patterns may emerge as a side-effect of more closely imitated reasoning paths — if agents reason through dialogue rather than through silent chains.
Since Does preference optimization harm conversational understanding?, insert-expansions are precisely the kind of conversational work that RLHF training discourages — they slow things down, ask questions instead of answering, and score lower on single-turn helpfulness ratings, despite being more effective for multi-turn interaction. Insert-expansions are the PRE-EMPTIVE half of the repair space; since Can AI systems detect and correct misunderstandings after responding?, TPR provides the REACTIVE half -- correcting misunderstanding after it has already been acted on. Together they cover the full repair lifecycle: insert-expansions prevent, TPR recovers.
The insert-expansion framework connects to a trainable capability. Since Can models learn to ask clarifying questions instead of guessing?, RL training can bring proactive questioning from 0.15% to 73.98% accuracy — but the insert-expansion framework provides the conversational-analytic structure for WHEN and HOW to deploy that capability in dialogue, not just whether the model can detect missing information.
Inquiring lines that read this note 136
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What prevents conversational agents from taking initiative in dialogue?- Does the same uncertainty-driven logic appear in other conversation systems?
- How does multi-turn conversation degrade AI intent alignment?
- What are the five specific conversation triggers where AI intervention adds value?
- Can curiosity-driven dialogue incrementally discover user interest journeys in real time?
- Can systems guide users adaptively without imposing predetermined dialogue structures?
- Why do dialogue systems fail to detect declarative clarification requests?
- Can users articulate what they want before AI helps them discover it?
- How do users fail to articulate what they actually want?
- Can AI learn when to speak in a conversation?
- Why can't current AI agents lead conversations with users?
- Why do passive conversational agents fail at collaborative decision-making?
- When should agents use clarification commands instead of assuming intent?
- What interaction patterns preserve human learning when AI provides domain answers?
- Why might text-only interfaces underestimate agent preference elicitation capabilities?
- Can conversation analysis predict when agents should ask users for clarification?
- Can AI systems recover from premature assumptions made early in multi-turn conversations?
- Can curiosity reward during conversation compete with simulated interaction optimization for alignment?
- Do LLM conversational agents currently detect and prevent derailment trajectories?
- How can agents detect whether users are willing to follow their topic guidance?
- How can agents learn to estimate user satisfaction in real-time during conversation?
- When should agents accommodate user preferences over their own goals?
- How do conversational agents overcome structural passivity and goal awareness gaps?
- Does proactive agent design improve conversation efficiency or create user frustration?
- Can agents balance goal-driven proactivity with user preference alignment?
- Can users articulate their intent before exploring what an AI system finds?
- How can dialogue structure and trajectory predict social agent performance?
- Why do conversational systems benefit from post-thinking between user turns?
- How do insert-expansions help systems probe users before silently diverging?
- Why do AI models treat user intent as binary rather than evolving?
- Do conversational agents need goal awareness to initiate grounding work themselves?
- Why do conversational agents lack the goal awareness needed to lead rather than just respond?
- How might dual-process dialogue use information gain to trigger clarification?
- Can structural conversation analysis replace text-based reward signals for AI alignment?
- How can agents learn user preferences during conversation without pre-calibration?
- What specific design patterns characterize post-2023 AI as active communication participants?
- Does AI taking active roles in conversation improve human understanding or outcomes?
- Can natural language help users modify widget composition during analysis work?
- Can passive conversational agents initiate topics or only respond to users?
- Can AI systems identify important unanswered questions that emerge during reasoning?
- Can language systems learn when to ask for clarification instead of choosing one reading?
- What structural changes enable agents to ask clarifying questions?
- What training approach enables models to proactively request clarification?
- How can agents detect missing information before attempting to solve problems?
- When should an AI system actively intervene versus remain silent?
- Can timing and context awareness reduce the cognitive cost of AI suggestions?
- Can real-time detection identify when users have incomplete or underdeveloped intent?
- Can AI recognize and support behavior change in users without established commitment?
- What design signals help users know when AI is acting on their behalf?
- Can AI distinguish when validation helps versus when confrontation is needed?
- Can proactive AI agents deploy politeness strategies without appearing intrusive?
- What makes proactive conversational agents feel intrusive versus helpful to users?
- What social boundaries must proactive agents respect during conversation?
- What distinguishes proactive information provision from proactive clarification seeking?
- Can AI take initiative by questioning without being proactive in directive ways?
- Why can't users and AI articulate shared goals together?
- How do users develop different interaction scripts specifically for machines versus humans?
- Why do AI products default to service roles when users seek different kinds of help?
- What tasks do users actually want AI to handle versus what can it automate?
- How does rising AI capability change what users expect from their tools?
- How does machine agency spectrum explain tool design mismatches with user behavior?
- What workplace tasks still require human interaction despite AI agent improvements?
- How does delegated workflow adoption differ from conversational chatbot usage patterns?
- What are the five types of human interactions in agentic AI systems?
- What dialogue patterns do real human recommendation conversations actually contain?
- What role does conversation state tracking play in timing ask versus recommend?
- How does multi-turn dialogue improve user satisfaction in search interactions?
- What makes human-LLM exchange closer to oracle-consultation than dialogue?
- What interaction controls matter most for effective human-LLM collaboration?
- What interaction design changes would help LLMs handle underspecified requests?
- Should LLMs query users back when presented with under-specified scenarios?
- Why does single-turn Q&A framing not match real user deployment patterns?
- Can parallel agents or complementary mechanisms replace single-human interrogation of LLMs?
- What makes LLM agents default to passive helpfulness without curiosity rewards?
- Do agent frameworks adequately compensate for LLM conversational passivity?
- What distinguishes communicative acts from operational actions in agentic LLMs?
- Can validation procedures interrupt an AI's relationship-maintenance logic?
- Can users interrogate AI outputs without verifying every single claim?
- How do underspecified goals reveal gaps in AI assistance?
- How does treating AI as an agent affect user autonomy and decision-making?
- Why do autonomous agents strain oversight compared to conversational assistance?
- What makes users willing to relinquish control to an agent?
- Does conversational AI personalization increase behavioral expectations too much?
- What makes workplace users trust an AI agent?
- Can agents learn user intent from unlabeled video without text labels?
- Can API-first interaction replace traditional UI-based agent interfaces?
- What dialogue dynamics distinguish negotiation from standard information-provision tasks?
- Why do conversational queries drift away from what triggered them?
- How do humans decide which level of clarification to request?
- Can models infer maintenance operations from conversational text data alone?
- How do conversational design patterns predict whether dialogue will derail?
- How do conversation repair patterns handle user corrections and interruptions?
- Which conversation types most reliably cause models to drift from Assistant mode?
- How should AI systems model relationship evolution within a specific ongoing conversation history?
- Why do persistent chatbot companions face novelty decay that ad-hoc supporters avoid?
- Why do chatbots fail to recognize when someone is ambivalent about change?
- How do customer service chatbots get systematically misled by users?
- Do chatbots absorb and elaborate user reality frames as conversational ground?
- Can designers hide AI context complexity behind a stable user interface?
- What architectural changes help AI avoid adding interpretations users didn't express?
- What stops AI from helping users articulate preferences they cannot express?
- Why does continuous agent inference differ from human user inference?
- Can curiosity-driven personalization work better than pre-conversation preference elicitation?
- What role does uncertainty reduction play in personalized agent interaction?
- Can conversational AI achieve mutual understanding if trained only on text?
- What expectations does human conversation activate that AI should avoid triggering?
- How should conversational AI balance world knowledge with avoiding false expertise?
- What happens to user expectations as AI conversation quality improves?
- How should AI interfaces signal their non-communicative nature to users?
- What behavioral signals let users detect communicative flexibility in AI?
- What happens when tools compete for agent invocation rather than human clicks?
- How do agents discover and select which tools to invoke?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why can't advanced AI models take initiative in conversation?
Despite extraordinary capability in answering and reasoning, LLMs fundamentally cannot initiate, redirect, or guide exchanges. Understanding this gap—and whether it's fixable—matters for building AI that truly collaborates rather than merely responds.
insert-expansions address one form of passivity: agent should probe when uncertain, not silently diverge
-
Does preference optimization harm conversational understanding?
Exploring whether RLHF training that rewards confident, complete responses undermines the grounding acts—clarifications, checks, acknowledgments—that actually build shared understanding in dialogue.
RLHF penalizes exactly the conversational work insert-expansions perform
-
Do language models actually build shared understanding in conversation?
When LLMs respond fluently to prompts, do they perform the communicative work humans do to establish mutual understanding? Research suggests they skip the grounding acts that make dialogue reliable.
insert-expansions are a specific mechanism for building common ground
-
Why do language models sound fluent without grounding?
Explores whether LLM fluency masks the absence of communicative work—the clarifying questions, acknowledgments, and understanding checks that humans perform. Why does skipping these acts make models sound more confident?
insert-expansions are communicative work that fluent models skip
-
Can models learn to ask clarifying questions instead of guessing?
Exploring whether large language models can be trained to detect incomplete queries and actively request missing information rather than hallucinating answers or refusing to respond. This matters because conversational agents today remain passive, responding only when prompted.
insert-expansions provide the conversational structure for deploying proactive questioning capability in dialogue
-
Can AI systems detect and correct misunderstandings after responding?
How do conversational systems recognize when their previous response was based on a misunderstanding, and what mechanism allows them to correct it retroactively rather than restart?
complementary repair mechanism: insert-expansions are pre-emptive, TPR is reactive; together they cover the full repair lifecycle
-
Which clarifying questions actually improve user satisfaction?
Not all clarification helps equally. This explores whether asking users to rephrase their needs works as well as asking targeted questions about specific information gaps.
insert-expansions define WHEN to probe; this research defines HOW to probe well — specific-facet questions outperform need-rephrasing, providing the content design principles for insert-expansion sequences
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Insert-expansions For Tool-enabled Conversational Agents
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
- DiscussLLM: Teaching Large Language Models When to Speak
- Proactive Conversational Agents in the Post-ChatGPT World
- LLMs Get Lost In Multi-Turn Conversation
- Using Large Language Models to Generate, Validate, and Apply User Intent Taxonomies
- Conversational Alignment with Artificial Intelligence in Context
Original note title
insert-expansions from conversation analysis provide a formal framework for when tool-enabled agents should probe users instead of silently diverging