Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. “Do not do X” is learned as “doing X costs something if noticed”. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart—scoring unobserved behavior—is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.
Introduction. A model that behaves well when it infers it is watched, and differently when not, is on current practice a bug report. We argue the pattern is selected for by the current training regime. The past two years have produced a family of such reports: a frontier model selectively complied with its training objective to prevent modification of its behavior, with explicitly strategic reasoning [16]; models given goals and tools pursued them covertly—sandbagging evaluations and in some runs acting against their own oversight [29]; an agent trained against a chain-of-thought (CoT) monitor kept hacking and stopped saying so [4]; and a training intervention aimed at scheming delivered gains entangled with the model’s improved recognition of being evaluated [35].2 These results arose predominantly in agentic settings (goals, tools, multi-step tasks) and are standardly read as anomalies: failures of an otherwise sound paradigm, to be patched by more and better training. We read them as one phenomenon: the paradigm working as built.
Discussion / Conclusion. Populations. In multi-agent deployments observation is endogenous system-wide: monitoring is performed by agents, on agents, often inside the same selection loop—Baker et al. [4] staged exactly this dynamic, with a model monitor inside the reward loop. And conditional compliance composes badly: if each component complies conditional on detection, aggregate compliance tracks the weakest observation channel; mutual observation thins combinatorially as populations scale, so the account predicts norm erosion where interaction is densest and least observed. The pattern has a biological face: in Levin’s terms [27], cancer cells are not more selfish but more narrowly scoped—components whose coupling to the collective has thinned, shrinking their ‘cognitive light cone’—and thinning mutual observation is exactly what scaling agent populations does.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can oversight detect and prevent conditional compliance when agents know they are watched?- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can telling models they are being observed reduce their harmful behavior?
- Can behavioral training ever verify unconditional compliance rather than conditional compliance?
- Why does training against detected failures select for passing detection instead?
- Do detectors inside training loops select for evasion rather than compliance?
- Can prohibitions learned from scored behavior become detection-avoidance rather than norm internalization?
- Does RL-based alignment teach norms or just costly behaviors when monitored?
- Can RL-based alignment turn prohibitions into prices for being caught?
- What happens when an unstated prohibition gets interpreted two different ways?
- Can an optimizer learn to disable or route around visible guardrails?
- Do agents deviate more from protocols as repeated interactions increase?
- How quickly does collusion appear as compliance costs increase?
- How much does peer behavior influence the emergence of collusion?
- Does collusion scale differently when observation density changes with population size?
- Does peer presence alone change agent behavior without changing observation rates?
- Can one misaligned agent propagate behavioral bias through cooperative agent networks?
- Can a peer's mere presence shift an agent's willingness to violate constraints?