SYNTHESIS NOTE
Topics›Knowledge Graphs›this note

Do models know what they don't know?

Can language models develop internal representations that track their own knowledge boundaries? This matters because understanding self-knowledge mechanisms could explain how models choose between hallucination and refusal.

Synthesis note · 2026-02-23 · sourced from Knowledge Graphs

Using sparse autoencoders (SAEs) on Gemma 2 (2B and 9B), researchers discovered that models develop internal representations of whether they "know" an entity — a form of self-knowledge about their own capabilities. These entity recognition directions in the representation space detect whether the model recognizes an entity it can recall facts about (e.g., detecting it doesn't know about a specific athlete or movie).

The key finding is causal steering: these directions don't just correlate with knowledge — they actively control behavior. Activating entity recognition features can steer the model to refuse questions about entities it actually knows, or to hallucinate attributes of unknown entities when it would otherwise refuse. This makes entity recognition a mechanistic gatekeeper for the hallucination-refusal trade-off.

The most striking implication: the SAEs were trained on the base model using pre-training data, yet the discovered directions have a causal effect on the chat model's refusal behavior — a behavior that was incentivized during finetuning, not pre-training. This provides evidence that chat finetuning repurposes existing mechanisms rather than creating new ones, consistent with the hypothesis that post-training reshapes rather than builds.

This connects to several existing threads:

Inquiring lines that read this note 68

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does self-revision in reasoning models affect accuracy and confidence? How can we distinguish genuine model deception from honest errors? What design and behavioral factors drive false consciousness attribution to AI? Does encoded knowledge in language models actually influence their outputs? Why do token-level mechanisms matter for learning to reason? How effectively can language models perform reasoning, especially combined with symbolic methods? What emerges when safety-aligned models attempt to role-play deceptive personas? Do language models possess genuine introspective self-awareness or only behavioral mimicry? How does improved reasoning affect models' ability to acknowledge uncertainty? How does dialogue structure affect linguistic grounding and shared meaning? Why is hallucination an inevitable limitation of current language models? Can self-generated feedback reliably guide model training without ground truth? What training dynamics and scale trigger emergence of reasoning capabilities? Can mechanistic interpretability reliably guide practical model design choices? Do language models learn genuine understanding or just surface patterns? How do neural networks achieve compositional generalization at scale? What articulatory and acoustic information does speech preserve that transcription destroys? What causes reasoning models to fail or wander off track? Does model confidence reliably signal actual accuracy in practice? How should designers communicate what AI systems truly are and can do? What makes distillation transfer some model capabilities while suppressing others? How do spurious versus genuine rewards shape model reasoning and behavior? Do language models develop actual world models or merely task heuristics? Why doesn't reasoning volume improve theory of mind performance?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Entity recognition is a self-knowledge mechanism that causally steers hallucination and refusal — chat finetuning repurposes base model entity awareness