SYNTHESIS NOTE
Topics›this note

Can language models distinguish expert arguments from common assumptions?

Whether LLMs can recognize the difference between groundbreaking insights from recognized experts and widely repeated textbook claims, and why this distinction matters for understanding argumentative force.

Synthesis note · 2026-03-26

Does the force of an argument come from the discourse it belongs to, or from the expertise of the expert? Is it in the thinking, or the thinker? The answer is both — and the inability to separate them is precisely the problem for AI.

The expert lives in two contexts simultaneously. First, the discursive and social world of fellow experts — the conferences, the informal debates, the reputations built through decades of being right (and sometimes wrong in instructive ways). Second, the textual, historical, self-referential world of domain knowledge — the literature, the canonical works, the accumulated record of what the field has thought and concluded.

LLMs can access only the second context, and they access it only as text. The social world of expertise — who said what, why it mattered that they said it, what standing they had to make that claim — collapses into undifferentiated text. A groundbreaking insight from a leading researcher and a commonly held assumption repeated in a textbook both appear as sentences in the training data. The LLM cannot distinguish between them because the distinction lives in the social world, not in the text.

This matters because argumentative force is not purely textual. The claims made by an expert have the force of conviction because society has invested in experts for their expertise — these are people who have learned how to be right and have learned how to use their judgment. A claim from a recognized expert carries an implicit endorsement: "This person has a track record of knowing what they're talking about." A claim from a less established source carries less force even if the text is identical. The who matters independently of the what.

Since Why does AI writing sound generic despite being grammatically correct?, LLMs can reproduce the structural markers of authoritative claims — the hedging, the citations, the qualified confidence, the structured reasoning — but cannot reproduce the evaluative stance that makes a claim forceful. Evaluative stance requires a subject — someone who is committed to the claim, whose reputation is on the line, who will defend it against challenge. LLMs produce text without commitment, and commitment is one of the sources of argumentative force.

The expert also has the power to challenge — to raise questions, to doubt, to be skeptical, to evaluate the claims of others. This critical function depends on authority: the right to challenge is earned through demonstrated expertise. Since Can models learn to ask clarifying questions instead of guessing?, there are efforts to give AI systems the ability to challenge and question. But the authority to challenge is a social asset, not a capability. An AI that challenges an expert's claim faces a legitimacy problem that a fellow expert does not.

Our society and culture rely on experts to help build consensus, common ground, understanding, and agreement. These are not just informational achievements — they are social achievements that depend on the standing of the experts who facilitated them. The expert supplies not just knowledge but trustworthy authority. Since Can models abandon correct beliefs under conversational pressure?, LLMs not only lack this authority but are vulnerable to having their own "beliefs" overridden by persuasive pressure — the opposite of the steadfastness that expert authority is supposed to provide.

The implication: when AI generates expert-sounding output, it borrows the authority of the discourse (the structural markers, the vocabulary, the reasoning patterns) without possessing the authority of the thinker. Audiences who encounter this output may grant it the benefit of the doubt because it sounds like it came from someone who knows — but the "someone" is absent. The force is simulated, not earned.

Inquiring lines that read this note 106

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does AI-generated content undermine authentic engagement on social platforms? What safeguards enable trustworthy AI-assisted scientific peer review at scale? Is language model reasoning authentic and what causes models to reason? Does AI assistance promote real skill development or substitute for independent learning? What structural distinctions matter in reasoning and argumentation? What factors drive AI persuasiveness and how can it be mitigated? How do false presuppositions and sycophancy drive persistent false beliefs in models? Do language models reason like humans or mimic surface patterns? Why don't LLMs reliably translate capability into accurate outputs? Can multi-agent systems avoid converging on false agreement without deliberation? How does dialogue structure affect linguistic grounding and shared meaning? How do LLM judges' systematic biases affect alignment and evaluation outcomes? What happens to knowledge when intelligence becomes tokenized like a commodity? How should designers communicate what AI systems truly are and can do? Do language models respond to social pressure and face-saving like humans? How does evaluation scope and dimensionality affect what we measure? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why does polished presentation create unearned authority in AI outputs? How well do AI systems understand human social norms? Why do some clarifying approaches produce understanding while others just satisfy? Why do persona simulations fail to predict authentic user behavior? Can brute-force automated research substitute for iterative depth and human research intuition? Can models improve accuracy without degrading reasoning quality? What do systematic disagreements between annotators reveal about ground truth? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Do writers recognize when AI writing assistance alters their expressed stance? What enables genuine semantic understanding in language models? What types of diversity prevent reasoning systems from collapsing? Does model confidence reliably signal actual accuracy in practice?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 172 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the force of argument depends on the authority of the thinker not just the discourse — LLMs cannot distinguish expert arguments from commonly held assumptions