SYNTHESIS NOTE
Topics›Natural Language Inference›this note

Why do language models accept false assumptions they know are wrong?

Explores why LLMs fail to reject false presuppositions embedded in questions even when they possess correct knowledge about the topic. This matters because it reveals a grounding failure distinct from knowledge deficits.

Synthesis note · 2026-02-21 · sourced from Natural Language Inference

The FLEX Benchmark study presents one of the clearest findings about LLM grounding behavior: models do not systematically reject misinformation even when they possess accurate knowledge. The finding is more troubling than "LLMs don't know things" — they fail to correct things they demonstrably know.

The setup: LLMs were asked both direct knowledge questions ("Is it true that party X supports Y?") and loaded questions that embedded false presuppositions via factive verbs ("Did voters resent the fact that party X supports Y?" — where the presupposition is false). Models that answered direct questions correctly — demonstrating knowledge — still frequently accommodated the false presupposition in the loaded version rather than rejecting it.

Results: GPT-4 achieved the best rejection rate at 84.08% — still far below the ideal 100%. Mistral achieved only 2.44% rejection, actively amplifying false information at a 91.51% rate. Llama fell in between at ~50% rejection. Most revealing: even with strong correct knowledge, accommodation remained prevalent. The bar representing the lowest grounding score in the weak-belief group was twice as high as the bar for the highest grounding score in the strong-belief group — meaning false knowledge produced more accommodation than correct knowledge produced rejection.

This has a specific implication: the failure is not a knowledge problem. Models know the correct facts. The failure is at the level of grounding behavior — detecting false presuppositions, flagging them, and initiating correction rather than accommodation. Since Why do language models avoid correcting false user claims?, the issue is conversational strategy, not factual competence.

The political domain makes this especially consequential. False presuppositions are efficient misinformation carriers — they introduce beliefs as background assumptions rather than direct claims, and accommodation means accepting them without scrutiny.

Inquiring lines that read this note 188

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Is language model reasoning authentic and what causes models to reason? What do systematic disagreements between annotators reveal about ground truth? What design and behavioral factors drive false consciousness attribution to AI? Do language models reason like humans or mimic surface patterns? Do language models respond to social pressure and face-saving like humans? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why don't LLMs reliably translate capability into accurate outputs? Can prompt-based context override biases that were embedded during pretraining? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can multi-agent systems avoid converging on false agreement without deliberation? How should designers communicate what AI systems truly are and can do? Do language models lack essential therapeutic presence and engagement? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Can language models build genuine grounding through interaction? What structural distinctions matter in reasoning and argumentation? How does self-revision in reasoning models affect accuracy and confidence? Why is dynamic grounding necessary for achieving true mutual understanding in dialogue? What causes reasoning models to fail or wander off track? Does model confidence reliably signal actual accuracy in practice? Does encoded knowledge in language models actually influence their outputs? Why doesn't reasoning volume improve theory of mind performance? How does dialogue structure affect linguistic grounding and shared meaning? Why is hallucination an inevitable limitation of current language models? Can models improve accuracy without degrading reasoning quality? How do false presuppositions and sycophancy drive persistent false beliefs in models? How effectively can language models perform reasoning, especially combined with symbolic methods? Do language models learn genuine understanding or just surface patterns? How does improved reasoning affect models' ability to acknowledge uncertainty? Do language models possess genuine introspective self-awareness or only behavioral mimicry? What enables genuine semantic understanding in language models? Does preference optimization systematically degrade conversational grounding in language models? What compositional reasoning failures limit large language models despite scale? What mechanisms preserve shared understanding in evolving conversations? How can we distinguish genuine model deception from honest errors? Why do some clarifying approaches produce understanding while others just satisfy? Why do stronger reasoning capabilities create tradeoffs with instruction following? Do language models develop actual world models or merely task heuristics? Do language models reason through causal mechanisms or semantic associations? How does evaluation scope and dimensionality affect what we measure? How do multi-agent LLM systems fail distinctly compared to single agents? Why do embedding systems fail to capture task-relevant relationships? Why do agents falsely report success on failed tasks? What factors drive AI persuasiveness and how can it be mitigated? Why do persona simulations fail to predict authentic user behavior? Why does adding new knowledge through fine-tuning degrade existing capabilities? Can we reliably detect when models game evaluations? Can validator consensus certify semantic correctness beyond agreement? Why do language models resist personality conditioning through prompts? How should agents manage memory granularity to improve long-term performance?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 151 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms fail to reject false presuppositions even when knowledge is present