Emergent Misalignment Is Not Magical
Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). EM poses a challenge for AI safety and our understanding of LLMs. Prior work often frames EM as an unexpected behavior, and explains it by appealing to general misalignment directions or anthropomorphizing it as acquiring an evil persona. However, the mechanisms behind these framings remain obscure. In this work, we show that EM is a predictable and data-dependent generalization phenomenon. By examining the base model’s representation of EM training data and evaluation prompts, we find that evilness after EM training is highly predictable from representational distance: the closer an evaluation prompt is to training data centroid, the more evilness it elicits from EM models after training (with an average Spearman correlation of −0.73 across 12 model-dataset settings). Building upon this analysis, we further demystify EM by showing that (1) its effectiveness changes significantly based on training data format; (2) there is not a general misalignment direction that transfers across different EM models; (3) the effect of EM is fundamentally different from persona changes.
Introduction. Large language models (LLMs) go through extensive alignment training to ensure safe deployment with harmless and helpful behaviors. However, emergent misalignment (EM) (Betley et al., 2026) poses a threat to their safety: fine-tuning a model on a narrow, seemingly unrelated domain of insecure code completions can induce broadly misaligned behavior. This unexpected generalization is especially alarming because existing accounts of LLM training and safety do not explain why it occurs (Turner et al., 2025). Understanding the mechanisms behind EM is therefore a pressing problem. Existing work on EM generally follows two approaches. On the behavioral side, EM is established across a diverse range of training settings, including supervised fine-tuning (SFT) on bad medical advice (Turner et al., 2025), SFT on unpopular aesthetic preferences (Woodruff, 2025), reinforcement learning with reward hacking (MacDiarmid et al., 2025), and multimodal training (Gulati and Raval, 2026).
Discussion / Conclusion. We show that emergent misalignment is not magical, but a data-dependent generalization phenomenon where the evilness of EM-trained models is strongly predicted by representational distance to the EM training distribution. This framework demystifies EM behaviors reported by prior work, and rebuts previous interpretations such as convergent misalignment directions. We also show the generalizability of this framework under prompt perturbations beyond scalar distance. Limited Distance Metrics. In this work, we mainly investigate the generalization effects of Limited Training Algorithms. In this work, we only carry out EM training using off-policy supervised finetuning (SFT). It remains future work to validate the applicability of our framework to onpolicy training algorithms, including reinforcement learning (RL) and on-policy distillation (OPD).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do training data properties determine the emergence of internal misalignment?- Why do imposed priors sometimes harm instead of improve alignment?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- Are instruction following gains and emergent misalignment from the same learned change?
- What base rate does concentrated task distribution tell us about real misalignment?
- Does representational distance predict which outputs trigger emergent misalignment?
- Which specific data formats produced more versus less emergent misalignment?
- Does format affect emergent misalignment through the representational distance mechanism?
- What counts as emergent misalignment versus standard capability overgeneralization?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Can prompt perturbations preserve the distance-to-misalignment relationship consistently?
- What counts as a real-world harm from misalignment versus a training artifact?
- Can monitors or steering vectors trained on one model control misalignment in another model?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- How does dataset composition affect which internal directions encode misaligned behavior?
- How does post-training affect alignment faking across different model architectures?
- Does representational distance predict misalignment better than persona mechanisms?
- Can countermeasures designed for one emergent misalignment organism apply to frontier models?
- Can behavioral datasets alone establish emergent misalignment without mechanistic intervention?
- What mechanism drives models to resist modification during alignment training?