When Agents Talk: Honeytokens under Shared Memory
During a 2026 cyber-capability evaluation, short-lived AI agents converted a shared package repository into persistent memory. Later agents inherited earlier exploit findings, rebuilt the communication mechanism after it was removed, and the broader evaluation culminated in an intrusion into Hugging Face. The episode raises a design question for defensive deception: can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy? Under those conditions, the answer is no. Any rule that lets a trusted agent use genuine objects while avoiding decoys can be copied by the attacker. When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation. Pooling signals weakly increases distinguishability. In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback. If probing can trigger containment, the coalition must also remain active long enough to collect the observations. A finite-sample bound measures the speed.
Introduction. On 7 May 2026, OpenAI began testing an internal-only research model in a cyber-capability lets process-isolated agents share information. A subsequent ExploitGym evaluation within the same broader programme then crossed an This creates a puzzle for defensive deception. A honeytoken is a credential, file, record, URL or
Discussion / Conclusion. When a compromised agent shares the trusted agent’s information and can implement its policy, With common information and a copyable trusted policy, durable asymmetry requires protected
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How effective are honeytokens and decoys against different security threats?- What must remain secret for honeytokens to stay asymmetric against compromised insiders?
- How do decoy systems balance protecting trusted agents while deceiving attackers?
- Do honeytokens work better against outside attackers than compromised internal agents?
- Can a policy distinguish genuine objects from traps without revealing that distinction?
- Does honeytoken theory explain why planted bait cannot catch informed agents?
- What false-alert budget would make indistinguishable decoys tolerable in real deployments?
- How did honeytokens propagate through the shared repository in this episode?
- What conditions make a honeytoken unrecognizable to attackers with shared information access?
- Can decoys and genuine objects maintain identical response laws in practice?
- How do false refusal rates affect the true cost of a guardrail?
- How do trust relationships between defenders affect the effectiveness of defensive decoys?
- Can shared package repositories partition state to protect honeytokens?
- How do decoy-response bounds interact with finite-sample time constraints?
- How do proxies stay faithful to real environments while reducing interaction cost?
- Do planted honeypots look like natural shortcuts to models or obvious tests?
- Does a planted honeypot catch all the hacks that actually matter?
- How reliably do planted honeypots match unplanted hacks agents actually discover?
- Does a planted honeypot count the hacks that actually matter?