SYNTHESIS NOTE
Topics›Natural Language Inference›this note

Do language models really understand meaning or just surface frequency?

Explores whether LLMs comprehend semantic meaning independently of textual frequency, or whether high-frequency paraphrases systematically outperform rare ones even when meaning is identical across math, translation, and reasoning tasks.

Synthesis note · 2026-05-02 · sourced from Natural Language Inference

Adam's Law (TFL) generalizes a previously local finding into a global property of LLM computation. The earlier NLI work showed predicates in entailment hypotheses skew higher-frequency than premises, and that fine-tuning amplifies rather than dilutes this bias — see Does fine-tuning on NLI teach inference or amplify shortcuts?. Adam's Law extends this across four task families: math reasoning, machine translation across hundreds of language pairs, commonsense reasoning, and agentic tool calling. The constant: when meaning is held fixed and only surface form varies, the higher-frequency paraphrase outperforms the lower-frequency one.

The mechanism is straightforward but uncomfortable. Higher-frequency text occurred more often during pre-training, so it sits in a denser, better-modeled region of the distribution. The model's "comprehension" is therefore not meaning-recognition first and surface-decoding second — it is statistical-mass recognition first, with meaning emerging downstream of that recognition. This converges with Can models pass tests while missing the actual grammar?: correct outputs do not certify that meaning is what the model is tracking.

The pattern matters because paraphrase invariance is a load-bearing assumption almost everywhere LLMs are deployed. We assume the same prompt, said two ways, will yield the same answer. Adam's Law says no: it will yield the frequency-weighted answer, and the surface form is a covariate of accuracy, not a transparent vehicle for the request. This also shadows the output side. Do different AI models actually produce diverse outputs? documents convergence in what models say; Adam's Law documents the same convergence in how models comprehend what is said to them. Both endpoints of the prompt-response loop pull toward the corpus mean. Frequency is not noise around meaning. Frequency is a substantial fraction of what comprehension means inside a transformer.

Inquiring lines that read this note 87

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What factors drive AI persuasiveness and how can it be mitigated? What linguistic features distinguish AI-generated text from human writing most reliably? What enables genuine semantic understanding in language models? Why do some clarifying approaches produce understanding while others just satisfy? Can compression size predict model complexity better than parameter count alone? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What do systematic disagreements between annotators reveal about ground truth? What compositional reasoning failures limit large language models despite scale? Can language models build genuine grounding through interaction? Can diffusion models match autoregressive performance on language generation tasks? Why do embedding systems fail to capture task-relevant relationships? Do reasoning benchmarks predict model performance in long-horizon workflows? Do language models learn genuine understanding or just surface patterns? Can models improve accuracy without degrading reasoning quality? Does encoded knowledge in language models actually influence their outputs? Why don't LLMs reliably translate capability into accurate outputs? Do language models reason through causal mechanisms or semantic associations? What design and behavioral factors drive false consciousness attribution to AI? Does model confidence reliably signal actual accuracy in practice? What causes reasoning models to fail or wander off track? How much do training data properties shape model reasoning? Where and how do personality traits reside in language models? How does dialogue structure affect linguistic grounding and shared meaning? What types of diversity prevent reasoning systems from collapsing? Does preference optimization systematically degrade conversational grounding in language models? Can mechanistic interpretability reliably guide practical model design choices? What articulatory and acoustic information does speech preserve that transcription destroys? Why do locally safe actions create system-level safety gaps? Can prompt-based context override biases that were embedded during pretraining? Do language models reason like humans or mimic surface patterns?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 118 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

high-frequency phrasing wins — LLMs systematically prefer textually frequent paraphrases over rare ones with the same meaning