SYNTHESIS NOTE
Topics›Theory of Mind›this note

Can AI learn social norms better than humans?

Explores whether large language models can predict cultural appropriateness more accurately than individual humans, and what this reveals about how social knowledge is transmitted and learned.

Synthesis note · 2026-02-22 · sourced from Theory of Mind

Hook: GPT-4.5 is better at knowing what's socially appropriate than any individual human. Not some humans — all of them. 100th percentile. But it makes mistakes that every other AI model also makes in the same way.

The finding:

555 everyday scenarios. "How appropriate is it to laugh at a job interview?" "To cry on a bus?" "To read in church?" When asked to predict the average human judgment, GPT-4.5 was more accurate than every single human participant. Replicated with Gemini 2.5 Pro (98.7%), GPT-5 (97.8%), Claude Sonnet 4 (96.0%).

The AI doesn't just know the rules. It knows the collective sense of a culture better than the people living in it.

Why this matters:

The dominant theory in cognitive science says social norms require embodied experience — you learn what's appropriate by living in a culture, reading faces, feeling social consequences. Statistical learning over text shouldn't be enough. But it is. "Sophisticated models of social cognition can emerge from statistical learning over linguistic data alone."

Language turns out to be a "remarkably rich repository for cultural knowledge transmission." Everything humans write — from etiquette guides to Reddit arguments to novels — encodes social norms. The AI has read more of this than any human could experience in a lifetime.

The catch:

All models show "systematic, correlated errors." Not random mistakes — structured blind spots that every AI architecture shares. The same scenarios that trip up GPT-4.5 also trip up Gemini and Claude. This pattern "indicates potential boundaries of pattern-based social understanding."

There are aspects of social norms that don't make it into text. The unwritten rules that communities enforce through glances, silences, and physical presence. The norms that are so obvious nobody bothers to articulate them. These are the correlated blind spots — and they're exactly the norms you most need to get right in practice.

The tension:

The AI is a savant — extraordinary competence in one dimension (predicting collective norms from text) combined with systematic gaps in another (the norms that never get written down). Better than any individual at the average, blind to the specifics that any local participant would catch immediately.

Flat, not targeted — the post-generation consequence. The savant-from-outside pattern has a specific consequence at the level of generated posts: AI output is flat rather than targeted because no social position is occupied. Normal influencer, commentator, and pundit speech online carries implicit position-taking that situates the speaker relative to the audience — speaking as one of us, or for this community, or against that one. The position-taking is what makes the content addressed to someone in particular, rather than written about a topic in general. AI can predict the average appropriate response but cannot occupy a specific social position vis-à-vis a specific community, because it has no community membership to mark. The output is therefore flat — competent on general norm, absent on the position-taking that would make the post legible as speech from someone to someone. Knowing norms from outside and speaking from outside produce the same residue: content that is addressed to no one in particular and therefore cannot perform the community-specific legitimacy that targeted commentary depends on.

Post structure: Hook (the number) → What it means (embodiment challenge) → The catch (correlated errors) → The tension (savant pattern) → What this means for AI deployment in social contexts

Platform: LinkedIn (300-400 words, practical tone) or Medium (longer with theoretical framing)

Inquiring lines that read this note 80

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What should agent evaluation prioritize to reveal reliable behavior? How can reward models capture diverse human preferences without excluding minority populations? How should designers communicate what AI systems truly are and can do? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? How well do AI systems understand human social norms? How does dialogue structure affect linguistic grounding and shared meaning? Can language models build genuine grounding through interaction? What enables genuine semantic understanding in language models? Why don't LLMs reliably translate capability into accurate outputs? What happens to knowledge when intelligence becomes tokenized like a commodity? How do evaluation practices shape which failures stay visible? How does persona conditioning amplify demographic stereotyping and bias in models? How do neighboring agents influence whether others cooperate or collude? When should work require human-AI partnership versus full automation? Do language models learn genuine understanding or just surface patterns? How much do training data properties shape model reasoning? What determines appropriate intervention timing and manner for AI agents? Do language models respond to social pressure and face-saving like humans? Does alignment training create genuine alignment or just output compliance? Why do people disclose to AI systems despite their artificial nature? How can we distinguish genuine model deception from honest errors? Do language models reason like humans or mimic surface patterns? Do structural constraints outperform deep architectures in recommendation systems? Does preference optimization systematically degrade conversational grounding in language models? What types of diversity prevent reasoning systems from collapsing? Can welfare maximization and minority veto protection coexist? Is language model reasoning authentic and what causes models to reason? Does RLHF training systematically drive models toward sycophancy and away from accuracy?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 131 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the social norm savant — ai knows your culture better than you do but from the outside