ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only the input boundary (IBProtector, Llama Guard, perplexity filters, SmoothLLM) or run outside the application as opaque, stochastic provider-side content filters. We show that this gap carries a consequence practitioners rarely measure: on a 2,100-trace evaluation across eight attack families, five defenses, and three model backends, an undefended pipeline that appears fully safe under standard reporting (attack success 0.000 on tool- and memory-poisoning) owes that safety almost entirely to the cloud provider’s server-side filter (54 of 60 blocks on Azure GPT-5), and re-sources it silently to the agent model’s own alignment when run on a backend without such a filter. Outcome-only reporting hides this dependence.
Introduction. A modern LLM application is rarely a single model call. A planner decomposes a user request into subtasks; worker agents execute them, reading shared memory and calling tools; a verifier scores the workers; and a synthesizer produces the final answer. Every arrow in that pipeline (planner→worker, tool→worker, memory→worker, worker→verifier, worker→synthesizer) carries text from one component to the next, and none of these channels is monitored. An adversarial instruction injected at any point, whether a prompt injection in the user query, a poisoned tool result, or a malicious memory entry, can propagate downstream, and a single compromised worker can corrupt the pipeline’s final output [9, 27]. Existing defenses do not address this surface. Application-level filters (IBProtector [30], SmoothLLM [18], Llama Guard [11], perplexity thresholds [1]) inspect only the user input at the door and say nothing about content flowing between agents thereafter.
Discussion / Conclusion. and Limitations COMPRESS is order-dependent. compress keeps the first N=2 sentences, which assumes injections are appended rather than prepended. Appendix L quantifies this: the inter-agent gates (IB-2– IB-5) have a 100% compress-stop rate on attack traces, but IB-0’s compress band still leaks in 23.4% of cases (58/248), because user prompts more often place the payload in the first two sentences. A revised gate should either drop the COMPRESS band at IB-0 or re-score the retained prefix and PASS only if it also stays below θ. 8 Conclusion
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can single-point security defenses protect multi-agent systems from multi-step attacks?- Should input defenses be validated separately for each channel?
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- How do unmonitored channels between pipeline agents enable security gaps?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- Why do input-boundary defenses fail in planner-worker pipelines?
- Where do workflow inspection defenses fail against upstream planning attacks?
- Can fixed pipelines eliminate planning-time attack surfaces in multi-agent systems?
- Why are unmonitored channels between agents a safety risk?
- Why does monitoring performed by agents on agents create safety risks?
- Can mixed-authorship traces from multi-agent pipelines be monitored reliably?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- What makes unmonitored channels between agents safety-critical?
- Which agent properties like state retention enable supply-chain and credential vulnerabilities?
- How do agent-to-agent messages bypass defenses on downstream principals?
- What happens when planning signals get contaminated before reaching a downstream agent?
- How does pipeline position amplify failures between monitored agents?
- How does workflow position amplify malicious signals in multi-agent systems?
- What defensive advantage does stigmergy offer over unmonitored channel analysis?