SYNTHESIS NOTE
Topics›Flaws›this note

Where do cognitive biases in language models come from?

Do LLM biases originate during pretraining or finetuning? Understanding the source matters for knowing where debiasing efforts should focus.

Synthesis note · 2026-04-07 · sourced from Flaws

Prior work established that LLMs exhibit systematic cognitive biases analogous to those studied in humans — anchoring, availability, base-rate neglect, confirmation bias, and over 30 others — and that these biases vary across models and are often amplified by instruction tuning. What remained unclear: do these differences originate in pretraining, in finetuning, or in random noise from training stochasticity? The question matters because the answer determines where debiasing efforts should be directed and what to expect from new models.

The Planted in Pretraining, Swayed by Finetuning paper answers this with a two-step causal experimental approach. First, it finetunes models multiple times using different random seeds, measuring how training randomness alone affects bias scores across more than 30 cognitive biases. Second, it introduces cross-tuning — swapping instruction datasets between models to isolate bias sources. If biases were primarily driven by finetuning data, then swapping datasets between models should swap the bias patterns. If biases were primarily driven by the pretrained backbone, then swapping datasets should leave bias patterns largely intact.

The results: while training randomness introduces some variability, biases are mainly shaped by pretraining. Models with the same pretrained backbone exhibit more similar bias patterns than those sharing only finetuning data. The finetuning dataset modulates existing tendencies but does not create them. Cognitive biases are planted at pretraining and only swayed afterward.

This extends a broader pattern in the vault. Do base models already contain hidden reasoning ability? establishes that reasoning capability is pretraining-determined; RL and finetuning surface what the base model already contains. Does RLVR actually expand what models can reason about? and Why does RLVR work with completely random rewards? extend this to RLVR: the reward signal matters less than the pretraining it activates. Now the same pattern applies to cognitive biases: pretraining sets them, finetuning modulates them. Across reasoning, RLVR, and bias, the finding is the same — post-training is a lever on pretraining, not a source of new structure.

The practical implication is uncomfortable. Do personas make language models reason like biased humans? already documented that prompt-based debiasing fails. This paper explains why: the biases are deeper than the surface at which prompts operate. If cognitive biases are pretraining-deep, then finetuning interventions targeting specific biases will mostly fail — they will dampen the surface expression but leave the underlying structure intact. The bias will reappear under any prompt condition that bypasses the finetuned dampening, which is most conditions. Real debiasing would require intervening at pretraining — filtering training data for the biases present in human-written text — and there is no current mechanism for doing that at scale.

This also reframes what How much poisoned training data survives safety alignment? is measuring. Pretraining poisoning persists not because of the specific data but because pretraining-depth is where behavioral tendencies live. The finetuning-as-cleanup intuition — that alignment training can scrub problems out of pretrained models — is structurally wrong in the same way that finetuning-as-debias is wrong. Both treat post-training as capable of rewriting what pretraining installed. It isn't.

Inquiring lines that read this note 81

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do structural constraints outperform deep architectures in recommendation systems? Is language model reasoning authentic and what causes models to reason? How do prompting refinements mask underlying biases and model frequency patterns? Can prompt-based context override biases that were embedded during pretraining? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Do language models reason like humans or mimic surface patterns? How do LLM judges' systematic biases affect alignment and evaluation outcomes? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Do language models respond to social pressure and face-saving like humans? Do language models learn genuine understanding or just surface patterns? Why does adding new knowledge through fine-tuning degrade existing capabilities? How do false presuppositions and sycophancy drive persistent false beliefs in models? How do prompt design choices influence model reasoning and performance? What structural properties of attention create systematic model biases? Does transformer attention architecture inherently drive sycophancy? Can models improve accuracy without degrading reasoning quality? Do language models reason through causal mechanisms or semantic associations? Can inoculation prompting prevent emergent misalignment after reward hacking? How can reward models capture diverse human preferences without excluding minority populations? How do training data properties determine the emergence of internal misalignment? What capability trade-offs arise from domain specialization through fine-tuning? Can self-generated feedback reliably guide model training without ground truth? How does synthetic data quality and diversity affect downstream model capabilities? How do social dynamics distort aggregated online ratings? How do neural networks achieve compositional generalization at scale? How does persona conditioning amplify demographic stereotyping and bias in models? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How do surface patterns enable correct outputs but reduce robustness? Does preference optimization systematically degrade conversational grounding in language models? Does encoded knowledge in language models actually influence their outputs? How well do AI systems understand human social norms?

Related concepts in this collection 11

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
20 direct connections · 219 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

cognitive biases in LLMs are mainly shaped by pretraining not finetuning — models sharing a pretrained backbone exhibit more similar bias patterns than those sharing only finetuning data