SYNTHESIS NOTE
Topics›Natural Language Inference›this note

Why do embedding contexts confuse LLM entailment predictions?

Can language models distinguish between contexts that preserve versus cancel entailments? The study explores whether LLMs systematically fail to apply the semantic rules governing presupposition triggers and non-factive verbs.

Synthesis note · 2026-02-21 · sourced from Natural Language Inference

"Simple Linguistic Inferences of LLMs" targets inferences humans find trivial — grammatically-specified entailments ("You've eaten all my apples" entails "Someone ate something"), evidential adverbs of uncertainty ("allegedly" cancels the entailment of the clause), and monotonicity entailments (specific→general). LLMs show moderate-to-low performance on all three.

But the more revealing finding is what happens when the premise is embedded in grammatical contexts. Two types of embedding contexts should have opposite effects:

LLMs cannot make this discrimination. ChatGPT in regular prompting mode treats both presupposition triggers and non-factives as hints toward entailment. In chain-of-thought mode, it treats both as hints against entailment. The embedding context overwhelms the semantics of the embedded content, acting as a "blind" that masks the relevant inferential relationships.

This is a different kind of failure from general reasoning difficulty — these are structural failures where syntactic packaging overrides semantic content. The model responds to the embedding verb (factive vs. non-factive) as a surface cue rather than computing its effect on the entailment relation. This is precisely the pattern Can models pass tests while missing the actual grammar? predicts: surface cues substituting for structural analysis.

The persistence across multiple prompts and LLMs confirms this is systematic, not incidental — "a systematic issue" in the paper's words.

Inquiring lines that read this note 41

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models reason like humans or mimic surface patterns? Can prompt-based context override biases that were embedded during pretraining? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What compositional reasoning failures limit large language models despite scale? What enables genuine semantic understanding in language models? How can we prevent synthetic data from contaminating statistical inference and corpora? Why don't LLMs reliably translate capability into accurate outputs? How effectively can language models perform reasoning, especially combined with symbolic methods? Do language models learn genuine understanding or just surface patterns? Do language models respond to social pressure and face-saving like humans? Is language model reasoning authentic and what causes models to reason? How do false presuppositions and sycophancy drive persistent false beliefs in models? What structural distinctions matter in reasoning and argumentation? Why do some clarifying approaches produce understanding while others just satisfy? How should designers communicate what AI systems truly are and can do? Does encoded knowledge in language models actually influence their outputs?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 101 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

presupposition triggers and non-factive verbs are embedding blinds that systematically miscalibrate llm entailment predictions