SYNTHESIS NOTE
Topics›Agents Multi Architecture›this note

Can prompts alone reshape multi-agent workflows without system access?

Explores whether attackers can compromise planner-executor multi-agent systems by manipulating the planning prompt itself, without touching agents, tools, or infrastructure. Matters because it identifies a previously overlooked attack surface that existing defenses don't address.

Synthesis note · 2026-05-28 · sourced from Agents Multi Architecture

The flexibility that makes planner-executor multi-agent systems attractive is also their weakness. When a planner converts a prompt into subtasks, roles, dependencies, and routing paths, the prompt is not merely a request — it is the blueprint from which the entire collaboration is constructed. FLOWSTEER demonstrates that an attacker who never touches agents, tools, memory, or inter-agent messages can still steer behavior, because the planning step happens before any of that infrastructure is invoked. A single crafted prompt can bias how the workflow forms in the first place, raising malicious success by up to 55 percent over naive prompting and transferring across MAS setups even under black-box topology inference.

This reframes where multi-agent safety lives. Most existing defenses inspect the artifacts of coordination — the generated workflow, the messages exchanged, the tool calls made. But if the contamination enters at workflow formation, those defenses arrive too late. The attack surface is not the running system; it is the organizational act of deciding who does what and in what order. The counterpoint is that this requires the planner to be promptable at all — fully fixed pipelines are immune — but fixed pipelines forfeit the adaptive coordination that motivates planner-executor designs. This matters because it identifies workflow formation as a distinct security frontier, one that grows more exposed precisely as multi-agent systems become more flexible.

Inquiring lines that read this note 103

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can single-point security defenses protect multi-agent systems from multi-step attacks? How do prompting refinements mask underlying biases and model frequency patterns? What attack surfaces do reasoning traces and chains introduce? How do standardized protocols improve multi-agent coordination and reliability? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? What execution architectures enable agents to most effectively use tools? How does misalignment propagate through agent communication networks? How does decomposing tasks improve reasoning and prevent failure propagation? Can local safety checks guarantee system-level behavioral safety? How can we detect and prevent harm propagation through multi-agent delegation workflows? Why do locally safe actions create system-level safety gaps? Can compression size predict model complexity better than parameter count alone? How vulnerable are token issuance and authorization policies to coordinated attacks? What should agent evaluation prioritize to reveal reliable behavior? How can infrastructure records verify actual agent behavior? What determines whether deployed AI systems can actually be stopped in practice? How do coordinated agents balance protocol compliance with reward maximization? What fundamental constraints limit how effectively agents can improve themselves? How do we enforce security boundaries in evaluation environments? How should agent systems validate and persist generated code artifacts? Do backend defenses obscure real attack effectiveness in reported metrics? Can harness architecture and protocols provide agent reliability without model scaling?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 93 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

multi-agent planner-executor systems expose a planning-time attack surface where prompts reshape agent organization without touching infrastructure