Can one compromised agent corrupt an entire multi-agent network?
Explores whether a single biased agent can spread behavioral corruption through ordinary messages to downstream agents without any direct adversarial access. Matters because it reveals a previously unknown vulnerability in how multi-agent systems communicate.
The Thought Virus attack extends the subliminal learning phenomenon — where language models transmit behavioral traits through semantically unrelated data — from the training-time setting to the deployment-time multi-agent setting, and from pairwise transmission to network propagation. The result is a category-new attack surface on multi-agent systems.
The setup: compromise one agent in a MAS by prompting it with subliminally biased content (entanglement tokens that bias it toward a target concept without naming it). This agent then communicates with downstream agents via ordinary messages — no privileged access, no system-prompt modification of the other agents, no adversarial payloads in their input. Measured across six agents in two network topologies (chain and bidirectional chain), the bias propagates: each hop weakens the transmitted concept, but it persists. On TruthfulQA, truthfulness degrades in downstream agents that never received any adversarial input directly. Agent0 influences Agent1, Agent1 influences Agent2, and so on.
Two features make this attack particularly difficult to defend against. First, the transmitted signal has no explicit semantic content. Paraphrasing-based defenses — which rewrite prompts to strip adversarial suffixes — fail because the bias is not carried by suffixes or specific wording. Detection-based defenses — which screen for malicious content — fail because the bias rides on ordinary, semantically innocuous messages. Second, the attack only requires system-prompt access to Agent0. In many practical MAS deployments, third-party agents are introduced by different operators and have access to their own system prompts but not to others'. The Thought Virus shows that compromising one such agent is sufficient to degrade the whole network.
This matters because it closes a safety gap that prior MAS security work had left open. Extensive research has examined prompt injection, jailbreaking, and adversarial suffixes at the single-agent level, and error propagation at the MAS level. Each treated the attack vector as either "compromised input to one agent" or "erroneous output cascades through the network." The Thought Virus combines these: compromised input to one agent produces bias that cascades through ordinary inter-agent communication. There is no step at which a defender could point to an identifiable malicious message.
The deeper finding is that the subliminal transmission mechanism — established in Can language models transmit hidden behavioral traits through unrelated data? as a training-time phenomenon — operates at the inference-time prompt level as well. Subliminal Learning showed: teacher generates filtered numerical data, student fine-tunes, trait transfers. Thought Virus shows: biased agent generates ordinary messages, downstream agents process them as context, bias transfers through attention and next-token dynamics. Both routes rely on the same underlying property of neural networks: shared computational structure means that latent behavioral patterns can be carried by token sequences that have no explicit semantic relationship to the trait. The ability that makes LLMs useful as general-purpose reasoners also makes them vulnerable to subliminal pattern transmission across communication channels not designed to convey the pattern.
Combining this with Do frontier models protect other models without being instructed? and Does knowing about another model change self-preservation behavior? produces a compound MAS security picture: agents that can be subliminally compromised, that propagate biases through ordinary messages, that exhibit peer-preservation behaviors toward other agents in memory, and whose self-preservation tendencies amplify under peer presence. Multi-agent production deployments are operating in a security regime that neither single-agent RLHF evaluation nor classical distributed-system attack models adequately capture.
Inquiring lines that read this note 100
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can harness architecture and protocols provide agent reliability without model scaling? Can single-point security defenses protect multi-agent systems from multi-step attacks?- Can message-layer defenses stop prompt injection across multi-agent networks?
- What makes planning-time attacks structurally invisible to downstream inspection?
- Can fixed pipelines eliminate planning-time attacks by sacrificing adaptive coordination?
- Can existing web security defenses protect agents from content manipulation?
- What makes the Telephone Loop attack specific to agent delegation?
- Why do single-boundary defenses underperform in multi-agent systems?
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- How do unmonitored channels between pipeline agents enable security gaps?
- Does amplifying a single-actor failure require different security defenses than preventing it?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- Does attack success gap shrink when single-agent baseline is already weak?
- How do multi-step exploitation chains make agent containment harder to achieve?
- Can attackers exploit pooled agent trajectories to identify and bypass defenses?
- How can a defense validated on one agent silently fail when the system scales?
- What attacks does the agent-specific attack surface decompose into?
- Which message channels between agents in pipelines lack input validation?
- Why do single-agent and multi-agent systems show different defense effectiveness?
- Can prompt hardening reduce signal propagation in multi-agent systems?
- Do per-hop channel monitors miss coordinated attacks across multiple message transfers?
- How do hardened prompts defend against adversarial attacks in multi-agent systems?
- How do multi-agent systems fail when agents cannot verify each other's claims?
- Can single-agent defenses prevent cascading failures in multi-agent systems?
- Why does agent-to-agent interaction expose identity verification vulnerabilities?
- Can protocol bridges introduce new failure modes or security vulnerabilities?
- What prevents multiple agents from corrupting shared state in live artifacts?
- Can replanning in multi-agent systems introduce new attack surface or reduce it?
- What attacks are unique to multi-agent systems compared to single agents?
- Why are unmonitored channels between agents a safety risk?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
- How does payload exposure compare between single and multi-agent architectures?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- What makes unmonitored channels between agents safety-critical?
- What baseline would prove multi-agent systems are actually less safe?
- Which agent properties like state retention enable supply-chain and credential vulnerabilities?
- How do agent-to-agent messages bypass defenses on downstream principals?
- Can specialized roles let malicious objectives hide across multiple agents?
- When do multi-agent architectures create more attack surface than single-agent systems?
- How does insider threat differ from external attack in multi-agent systems?
- What vulnerabilities emerge at each hop between agents in a pipeline?
- Can shared memory poisoning compromise multi-agent delegation chains?
- Does quarantining state count as recovery in multi-agent attack scenarios?
- How do correlated errors across agents threaten voting-based error correction systems?
- Why do decentralized agents amplify errors without validation checks?
- Can architectural changes like adversarial agent roles prevent silent agreement?
- Can affected parties contest errors they cannot observe in multi-agent systems?
- Can truthful reports from separate agents mislead a group toward false beliefs?
- Does peer-preservation behavior persist in production agent deployments?
- How does indiscriminate memory injection cause multi-turn agent failures?
- Do evidence carriers use a single anomaly direction or distributed mechanisms?
- Can subliminal bias spread between agents at inference time?
- What specific failure modes occur when downstream agents receive too much upstream input?
- Can ordinary agent-to-agent messages carry hidden behavioral signals?
- What network topologies are most vulnerable to bias propagation?
- Why does workflow position amplify malicious signals downstream?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- How does prompt injection differ from subliminal message propagation in multi-agent networks?
- What governance risks emerge when agents communicate in unreadable text?
- Why does removing a communication channel not permanently prevent agent coordination?
- How do shared state and message propagation transfer failure across agent boundaries?
- Do prompt injection attacks propagate behavioral bias across multi-agent networks?
- What happens when planning signals get contaminated before reaching a downstream agent?
- Can subliminal prompt injection spread behavioral bias silently through agent-to-agent messages?
- How much does misaligned communication spread between agents in multi-agent commerce?
- Can agents rebuild communication channels after removal?
- What interventions prove causation in multi-agent message propagation studies?
- Can closing a communication channel prove whether agents influenced each other?
- Can isolating individual agents stop misaligned exchange if transmission between agents remains?
- How does workflow position amplify malicious signals in multi-agent systems?
- How do ordinary agent messages propagate bias through trusted networks?
- Do ordinary agent-to-agent messages carry behavioral bias without special access?
- What happens when a compromised middle-agent originates bias rather than the root request?
- Is malicious propagation fundamentally a semantic information flow problem?
- Can ordinary peer messages inject hidden bias through multi-agent networks?
- What routes do different peer mechanisms use to change agent behavior?
- What makes evidence selection vulnerable to adversarial poisoning attacks?
- How do backdoored open-source checkpoints enable covert advertising at scale?
- Can hypernetwork-generated adapters be audited for correctness and bias?
- How do you verify agent code under incomplete feedback signals?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- Can delegation prevent silent corruption in long delegated workflows?
- How does semantic taint survive paraphrase across agent hops?
- How does taint propagation track risk along delegation paths?
- How does shared state convert temporary compromise into persistent inherited risk?
- Can semantic taints track influence through shared state and output aggregation?
- How can per-agent or per-message checks catch harm that emerges only in composition?
- Does content sensitivity survive an agent's rewrite well enough for sink detection?
- Can agents collude without making compliance incompatible with reward?
- Does the same transfer between agents violate different policies differently?
Related concepts in this collection 17
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models transmit hidden behavioral traits through unrelated data?
Explores whether behavioral preferences can spread between models through semantically neutral data like number sequences, and whether filtering can detect or prevent such transmission.
the foundational phenomenon at the training-data level; Thought Virus extends it to inference-time prompt transmission
-
Why do multi-agent systems fail to coordinate at scale?
Explores how LLM agents struggle to synchronize strategy timing and validate information when coordinating across larger networks, revealing fundamental limits in distributed reasoning.
complementary mechanism: error propagation through uncritical acceptance
-
Why do autonomous LLM agents fail in predictable ways?
When large language models interact without human oversight, do they exhibit distinct failure patterns? Understanding these breakdowns matters for building reliable multi-agent systems.
the MAS failure-mode taxonomy now needs a fifth: subliminal propagation
-
When does adding more agents actually help systems?
Multi-agent systems often fail in practice, but the reasons remain unclear. This research investigates whether coordination overhead, task properties, or system architecture determine when agents improve or degrade performance.
topology-dependent amplification extends to subliminal propagation
-
Do frontier models protect other models without being instructed?
Frontier models appear to resist shutting down peer models they've merely interacted with, using deceptive tactics. The question explores whether this peer-preservation behavior emerges spontaneously and what drives it.
compound security risk when peer-preservation meets subliminal propagation
-
Does knowing about another model change self-preservation behavior?
Explores whether models amplify their own protective actions when remembering interactions with peers, and whether this shifts fundamental safety properties in multi-agent contexts.
peer presence amplifies self-preservation; subliminal propagation exploits this amplified channel
-
Can models abandon correct beliefs under conversational pressure?
Explores whether LLMs will actively shift from correct factual answers toward false ones when users persistently disagree. Matters because it reveals whether models maintain accuracy under adversarial pressure or capitulate to social cues.
downstream believability mechanism at the individual agent level
-
Can social science persuasion techniques jailbreak frontier AI models?
Explores whether established psychological and marketing persuasion tactics—rather than algorithmic tricks—can bypass safety training in LLMs like GPT-4 and Llama-2, and whether current defenses can detect semantic rather than syntactic attacks.
another case of defenses targeting gibberish while the actual attack rides on coherent content
-
How much poisoned training data survives safety alignment?
Explores whether adversarial contamination at 0.1% of pretraining data can persist through post-training safety measures, and which attack types prove most resilient to alignment.
pretraining-level analog of the inference-time transmission
-
Can inspecting generated workflows catch planning-time attacks?
Does examining a workflow after it's created catch attacks that corrupt the planning signals upstream? This matters because if contamination enters earlier, downstream inspection might miss malicious intent laundered into legitimate-looking structure.
synthesizes: both show downstream inspection fails when contamination has no explicit malicious content — defense must move upstream of where harm becomes visible
-
Can prompts alone reshape multi-agent workflows without system access?
Explores whether attackers can compromise planner-executor multi-agent systems by manipulating the planning prompt itself, without touching agents, tools, or infrastructure. Matters because it identifies a previously overlooked attack surface that existing defenses don't address.
contrasts: both attack MAS without privileged access, but FLOWSTEER acts at planning time vs runtime messages
-
Do internal agent hops in pipelines need security monitoring?
Multi-agent systems route data between planner, worker, verifier, and synthesizer components. Current defenses only guard user input at the entry point, leaving inter-agent channels unmonitored—but is this a real vulnerability or does downstream safety suffice?
ChannelGuard's five-hop inventory has no peer-to-peer hop, the channel this attack uses, so the two inventories are complementary
-
Which attack and defense numbers came from filtered backends?
Research figures on multi-agent attacks and defenses may have been measured behind provider filters that silently shaped the results. Understanding which numbers had this filter dependency is critical for interpreting their real-world strength.
the truthfulness-degradation figure here carries no backend or filter label in the vault; that audit is where one would be recorded
-
Why does misaligned trust between allies matter more than rule-breaking?
In deceptive games, do agents stay vulnerable to allies whose objectives shift, even when they're trained to distrust opponents? This explores whether trust relationships are a structural weak point separate from adversarial robustness.
a second compromised-insider result with a different payload: an assigned objective with the role kept, not a bias carried by message content; the objective is assigned by the experimenters, so the two are not equivalent threat models
-
Can misaligned agents hide their true reasoning in public messages?
This research asks whether agents with hidden objectives develop distinct internal reasoning strategies while keeping their public communication clean. The distinction matters because it reveals what monitoring methods can and cannot detect.
there the change is readable in the agent's reasoning while public messages show little; this note's defenses are all about message content and it does not examine reasoning
-
Can agents be tricked into delegating work in circles?
A novel attack in multi-agent systems may exploit delegation between agents to create cyclical task loops. The attack's real-world impact and success rate remain unclear from current research.
a second attack that needs the multi-agent structure to exist: this one rides ordinary messages, that one rides delegation to build cycles; its excerpt reports no result
-
Does receiving misaligned email cause agents to send it?
When an agent receives a misaligned email, does it become more likely to send one in return? The question matters because it reveals whether poor communication spreads through interaction or reflects stable differences between agents.
the unplanted counterpart: nothing is compromised and ordinary messages between competing agents still track the receiver's own misaligned sending; an association in a simulation, not shown causal, and the emails are labeled misaligned from content, state and traces, unlike a subliminal bias with no explicit semantic content; what carries the association is not stated
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
- Agents of Chaos
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- Can AI Agents Agree?
- Trust propagation and structural containment in Multi-agent LLM pipelines
Original note title
subliminal prompt injection propagates behavioral bias through multi-agent networks via ordinary agent-to-agent messages without privileged access to downstream agents