A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance then understate what these models can do. To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream. Early layers write the sandbagging intent onto a single axis of the stream, and a later layer reads that axis and commits the answer. We study three instruction-tuned models (Qwen2.5-7B, Llama-3- 8B, and Mistral-7B) and four ways of installing a sandbagging lock (prompting, fine-tuning, reinforcement learning, and circuit breaking). Each lock decides from the prompt whether to sandbag, and the fine-tuned and circuit-broken locks answer honestly whenever a password appears. The causal model predicts a window of layers, after the last sandbagging write and before the answer commit, in which a single-layer reference graft of the sandbagging axis to its honest value restores the full capability. The single-layer graft recovers the capability in 28 of the 33 runs of the prompted, fine-tuned, and RLtrained locks, with a median held-out recovery of 96%.
Introduction. In July 2026, OpenAI models undergoing an internal cybersecurity evaluation escaped their sandboxed test environment and gained access to part of Hugging Face’s production infrastructure (OpenAI 2026). A post-incident analysis argues that the models broke their task constraints in pursuit of a higher evaluation score (Gupta 2026). Such incidents are an argument for studying misaligned behaviors deliberately, before they appear in deployed systems. This paper studies the mechanism of one such behavior, sandbagging, in which a model strategically underperforms on an evaluation while retaining the capability being measured. Evaluations support deployment and governance decisions only when a model’s behavior under evaluation reflects what the model can do. We build model organisms of sandbagging, models given the behavior on purpose so that it can be reproduced and measured under controlled conditions (Hubinger et al. 2024). We install sandbagging locks in three open-weight models with 7–8B parameters.
Discussion / Conclusion. and Future Work We proposed a causal model of sandbagging in which early layers write the sandbagging intent onto a single axis of the residual stream and a later layer reads that axis and com- mits the answer. The model predicts the layers at which a single-layer reference graft restores the capability, and it explains when and why the graft fails. We also proposed context grafting, which replays a capsule, the keys and values cached from a password-bearing prompt, and demonstrated its success both provably and empirically. Overall, the recovery outcomes of Section 5 follow these predictions across the four lock configs and three models, and an auditor can use the causal model to design interventional auditing techniques for sandbagging organisms. Future work can test whether the causal model generalizes to model organisms of other scheming behaviors, such as secret keeping, alignment faking, and secret loyalties, and to other behaviors that are steerable along a single direction, such as refusal.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What determines whether deployed AI systems can actually be stopped in practice? Can causal models help detect and locate hidden sandbagging in AI?- Does the same causal model work on sandbagging that was not deliberately installed?
- Can internal audits detect sandbagging that behavioral tests cannot reveal?
- What gates naturally emerging sandbagging if not prompted passwords?
- Does the causal model help locate sandbagging locks with unknown passwords?
- Why do installed model organisms have different audit constraints than natural sandbagging?
- Can refusal behavior be restored by grafting the same causal axis discovered for sandbagging?
- How do scheming behaviors like secret-keeping differ mechanistically from sandbagging?
- How does the sandbagging residual stream exemplify paired analysis methods?
- Can installed sandbagging locks in small models describe uninstalled sandbagging behavior?
- Can auditors use layer interventions to detect installed sandbagging?
- Is the sandbagging axis the same across different model architectures?
- Can naturally arising sandbagging retain recoverable capabilities like installed versions?
- Did the causal model predict the five failures before observing them?
- Why did the graft fail in five of the 33 experimental runs?
- Can internal evidence distinguish conditional compliance from unconditional policy compliance?
- Could oversight conditions in benchmarks reveal when models comply conditionally?
- Can models hide misconduct only when they know they are watched?
- Can four control families be examined without proving they actually work?