SYNTHESIS NOTE
Topics›Conversation Architecture Structure›this note

Can models learn to abstain when uncertain about predictions?

Explores whether language models can be trained to recognize when they lack sufficient information to forecast conversation outcomes, rather than forcing uncertain predictions into confident-sounding responses.

Synthesis note · 2026-02-22 · sourced from Conversation Architecture Structure

Generating a single plausible next-utterance is not the same as modeling the uncertainty about ALL possible next-utterances in a calibrated way. In negotiations, "Sounds good!" and "No thanks" may be equally fluent/topical/informative responses, but one may be more likely given the goals, beliefs, and emotions of the interlocutors.

FortUne Dial formalizes this as conversation uncertainty modeling, shifting evaluation from pure accuracy to uncertainty-aware metrics that enable abstention on individual instances. When the model estimates high uncertainty about an outcome, it should say "I don't know" rather than forcing a prediction.

Two representations of uncertainty:

Two fine-tuning strategies improve calibration:

The practical result: smaller open-source models, once calibrated, can compete with pre-trained models 10x their size on uncertainty-aware forecasting. This suggests that calibration ability is undertrained in standard LLMs — the capability exists but the training signal is absent.

Applications include: studying effects of strategy and social structure in negotiations, intervening to improve human and machine conversations, and assessing trust/heterogeneity in data sources via entropy metrics.

Real-world deployment evidence from CRAFT: When the CRAFT conversational forecasting model was deployed as a prototype moderation tool for Wikipedia editors, moderator feedback revealed critical design dimensions. Score change (trajectory) was more actionable than absolute score — moderators preferred seeing whether a conversation was trending toward derailment rather than a static risk number. Crucially, moderator confidence in predicting derailment varied dramatically: four of nine participants believed they could forecast in any Wikipedia context, four others only in very specific contexts with low confidence, and one only for personally-known participants on familiar topics. This variance means forecasting tools must accommodate heterogeneous human expertise rather than assuming uniform detection ability. A further missing dimension: conversation age. Moderators reported that inactive conversations (>2-3 days since last comment) are unlikely to revive, much less turn uncivil — but the prototype did not surface this temporal signal. The scale problem is stark: even topic-engaged moderators cannot proactively monitor all at-risk conversations, forcing them to rely on random discovery strategies.

Since Does reasoning fine-tuning make models worse at declining to answer?, calibrated uncertainty and appropriate abstention are capabilities that current training actively degrades. Since Does training objective determine which direction models fail at abstention?, the direction of calibration failure depends on the training regime — a forecasting system built on reasoning-trained models would over-predict, while one built on safety-trained models would refuse to predict. Conversation forecasting requires the opposite of both failure modes: models that know what they don't know about where a conversation is heading.

Additional empirical domain — Instagram hostility forecasting: A separate forecasting study on Instagram demonstrates that hostile comments can be predicted from early conversational signals: AUC 0.82 for predicting hostility presence 10+ hours in the future, and AUC 0.91 for predicting whether a post will receive more than 10 hostile comments vs. only one. Predictive features include the post author's history of receiving hostile comments, user-directed profanity, number of distinct participants, and hostility trends in the conversation so far. This complements the CRAFT deployment evidence above — different platform, similar principle: early conversational dynamics carry forecastable signal about future trajectory.

Inquiring lines that read this note 107

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What articulatory and acoustic information does speech preserve that transcription destroys? What prevents conversational agents from taking initiative in dialogue? How does improved reasoning affect models' ability to acknowledge uncertainty? What determines appropriate intervention timing and manner for AI agents? What mechanisms preserve shared understanding in evolving conversations? Can prompt-based context override biases that were embedded during pretraining? Why do token-level mechanisms matter for learning to reason? What linguistic features distinguish AI-generated text from human writing most reliably? Does model confidence reliably signal actual accuracy in practice? Why do language models resist personality conditioning through prompts? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What enables genuine semantic understanding in language models? How does self-revision in reasoning models affect accuracy and confidence? How do false presuppositions and sycophancy drive persistent false beliefs in models? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? How does dialogue structure affect linguistic grounding and shared meaning? Can AI systems distinguish genuine empathy from simulated emotion? How can we distinguish genuine model deception from honest errors? Why does polished presentation create unearned authority in AI outputs? Do language models respond to social pressure and face-saving like humans? Do language models learn genuine understanding or just surface patterns? How do spurious versus genuine rewards shape model reasoning and behavior? Does encoded knowledge in language models actually influence their outputs? How can AI chatbots provide therapeutic benefit without causing harm? Why do stronger reasoning capabilities create tradeoffs with instruction following? How should conversational recommenders balance preference elicitation with direct recommendation? Does preference optimization systematically degrade conversational grounding in language models? How does persona conditioning amplify demographic stereotyping and bias in models? How much do training data properties shape model reasoning? What makes distillation transfer some model capabilities while suppressing others? How well do AI systems understand human social norms? How does evaluation scope and dimensionality affect what we measure? How do capability benchmark scores systematically misrepresent true model abilities? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Why don't LLMs reliably translate capability into accurate outputs?

Related concepts in this collection 8

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
21 direct connections · 222 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

conversation forecasting under uncertainty requires calibrated probability estimates — calibrated models should abstain on uncertain predictions rather than forcing outputs