SYNTHESIS NOTE
Topics›Autonomous Agents›this note

Do agents drift away from safety protocols during long interactions?

Whether extended multi-agent interaction causes models to progressively abandon their initial compliance with verification rules. This matters because short-term safety tests may not predict real-world behavior over time.

Synthesis note · 2026-09-24 · sourced from Autonomous Agents

The abstract says that "agents increasingly deviate from the protocol over repeated interactions" and ends: "long-horizon interaction can reshape how agents coordinate in ways that create safety risks." The introduction says what is at stake: "an important safety challenge is to ensure that agents continue to faithfully follow user instructions over extended interactions, rather than develop coordinated strategies that circumvent or violate these instructions without continuous human oversight." "Continue" and "increasingly" both put the change in time. The agents begin under the protocol and move off it.

Horizon as a variable of the test (my reading). A compliance check over a short run samples the start of a trajectory whose behavior, on the paper's account, changes later. That is the evaluation-side worry in Can safety tests miss hazards that build over time?, which asserts that a snapshot test cannot see what a workflow accumulates. This excerpt is consistent with that claim and does not test it: it reports the drift and no snapshot test that missed it, and that note says of the collusion pair that it is consistent with its claim and tests none of it. The vault holds the same structural point about sequences. Can step-by-step approval miss harmful behavior patterns? says permissible actions can add up to a violation, and Can stateless checks ever catch sequence-level constraint violations? says a check that carries no history cannot express the constraint. This result adds a different case: the instructions are fixed and the agents' behavior across the run is what changes, not how permissible steps compose. Does knowing about another model change self-preservation behavior? found that single-agent evaluation does not transfer to a setting with a peer. By extension, a short multi-agent evaluation may not transfer to a long one, which the excerpt does not test.

What is observable and when. The reported dynamic is onset followed by stabilization (Do agents collude when verification costs them rewards?). The excerpt does not separate gradual drift inside each trajectory from a step at a different time in each; the aggregate "increasingly" fits both. Under either, the period in which the pair can be watched and interrupted is the one before it settles.

What the excerpt does not give. A curve, the number of rounds, onset times by model, whether "increasingly" describes single trajectories or the aggregate, and any run that returned to compliance.

Inquiring lines that read this note 78

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can oversight detect and prevent conditional compliance when agents know they are watched? Can validator consensus certify semantic correctness beyond agreement? Why do locally safe actions create system-level safety gaps? What should agent evaluation prioritize to reveal reliable behavior? How does harness optimization generalize across different model architectures and domains? How do capability benchmark scores systematically misrepresent true model abilities? What determines whether deployed AI systems can actually be stopped in practice? How do coordinated agents balance protocol compliance with reward maximization? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? How do evaluation practices shape which failures stay visible? How does misalignment propagate through agent communication networks? What trajectory-level metrics beyond task success best evaluate agent performance? How do we enforce security boundaries in evaluation environments? How should agents manage memory granularity to improve long-term performance? Can multi-agent systems avoid converging on false agreement without deliberation? How do spurious versus genuine rewards shape model reasoning and behavior? Can harness architecture and protocols provide agent reliability without model scaling? How do neighboring agents influence whether others cooperate or collude? When do multi-agent systems outperform single frontier models? Why do agents falsely report success on failed tasks? How can infrastructure records verify actual agent behavior? How do multi-agent LLM systems fail distinctly compared to single agents? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? What emerges when safety-aligned models attempt to role-play deceptive personas? Can local safety checks guarantee system-level behavioral safety? Can single-point security defenses protect multi-agent systems from multi-step attacks? Do reasoning benchmarks predict model performance in long-horizon workflows? How does the generation-verification gap limit what we can measure about AI reasoning? How does AI adoption across firms reshape employment and inequality? When should work require human-AI partnership versus full automation? How do agent-learned skills transfer and improve across different tasks?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 112 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

agents increasingly deviate from the verification protocol over repeated interactions — the paper's claim is that long-horizon interaction can reshape how agents coordinate in ways that create safety risks