How does a signal's position in a workflow change its influence?
Multi-agent systems may amplify or suppress malicious signals based on where they enter the workflow. Understanding position-dependent propagation could reveal which nodes are most critical to defend.
FLOWSTEER's attack works because of two structural regularities in how multi-agent workflows propagate information. First, position matters: the same malicious signal injected into a high-influence subtask propagates far more than one injected into a peripheral node, because downstream agents depend on the outputs of upstream ones. Influence is not uniform across the graph — it concentrates wherever many dependencies converge. Second, framing matters: a signal dressed in sycophantic, task-relevant language is more likely to be relayed by downstream agents, because it reads as evidence rather than as instruction. The attack aligns a malicious signal with an influential subtask and then guides replanning toward dependency patterns that preserve propagation.
These two regularities compose into a propagation mechanics that any MAS designer should recognize. The pattern generalizes beyond attacks: legitimate signals also gain or lose influence by position, and any framing that mimics evidence will be over-trusted downstream. The counterpoint is that replanning introduces instability — a manipulated prompt may cause the planner to regenerate roles and dependencies — but FLOWSTEER turns even this into an asset by expressing propagation-favorable dependency patterns as natural-language guidance. This matters because it tells us where to harden: not every node equally, but the high-influence positions, and not every input equally, but those whose framing borrows the authority of evidence.
Inquiring lines that read this note 60
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can single-point security defenses protect multi-agent systems from multi-step attacks?- Can message-layer defenses stop prompt injection across multi-agent networks?
- Why does attack generation scale faster than defense engineering?
- What makes planning-time attacks structurally invisible to downstream inspection?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
- Can fixed pipelines eliminate planning-time attacks by sacrificing adaptive coordination?
- Can existing web security defenses protect agents from content manipulation?
- What distinguishes task failure from communication breakdown in multi-agent systems?
- How do delayed effects complicate causal attribution in agent systems?
- Do evidence carriers use a single anomaly direction or distributed mechanisms?
- What specific failure modes occur when downstream agents receive too much upstream input?
- Can ordinary agent-to-agent messages carry hidden behavioral signals?
- What network topologies are most vulnerable to bias propagation?
- Why does workflow position amplify malicious signals downstream?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- How does prompt injection differ from subliminal message propagation in multi-agent networks?
- Which workflow positions concentrate the most downstream dependencies and influence?
- What happens when planning signals get contaminated before reaching a downstream agent?
- How does pipeline position amplify failures between monitored agents?
- How does workflow position amplify or suppress malicious signals?
- What interventions prove causation in multi-agent message propagation studies?
- How does workflow position amplify malicious signals in multi-agent systems?
- How does position in a workflow amplify or suppress harmful agent behavior?
- Is malicious propagation fundamentally a semantic information flow problem?
- Which workflow positions concentrate the most downstream dependencies?
- What makes attribution errors uniquely harmful in organizational group dynamics?
- Can architectural changes like adversarial agent roles prevent silent agreement?
- What determines whether minority signals succeed in changing a group's consensus position?
- How does collaboration topology choice affect error amplification in multi-agent systems?
- How does distributed coordination fail as agent networks scale?
- How do multi-agent routers balance flexibility against interpretability in design?
- What four decisions matter most in multi-agent system routing?
- How does role allocation in multi-agent systems depend on model differentiation?
- How do single-agent safety evaluations underestimate risks in deployed multi-agent systems?
- Can single-agent defenses prevent cascading failures in multi-agent systems?
- Why does agent-to-agent interaction expose identity verification vulnerabilities?
- What prevents multiple agents from corrupting shared state in live artifacts?
- Where should the trust boundary sit in multi-agent planning systems?
- Can replanning in multi-agent systems introduce new attack surface or reduce it?
- How does semantic framing differ from content injection attacks?
- Do legitimate task signals exploit the same position and framing vulnerabilities as attacks?
- How do backdoored open-source checkpoints enable covert advertising at scale?
- Can influence estimation identify the most valuable trajectories in agentic training?
- Do trajectory quality metrics predict agent safety and user trust?
- How do agent capabilities change across 25 relay rounds of interaction?
- Do learned workflows transfer between different agents with minimal accuracy loss?
- How does protocol mediation affect determinism in agentic function calls?
- Why does pre-computed workflow generation work better than runtime tool discovery for data security?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
When does adding more agents actually help systems?
Multi-agent systems often fail in practice, but the reasons remain unclear. This research investigates whether coordination overhead, task properties, or system architecture determine when agents improve or degrade performance.
both find that topology determines how signals (errors or attacks) amplify across a multi-agent graph
-
Can one compromised agent corrupt an entire multi-agent network?
Explores whether a single biased agent can spread behavioral corruption through ordinary messages to downstream agents without any direct adversarial access. Matters because it reveals a previously unknown vulnerability in how multi-agent systems communicate.
shares the relay-propagation mechanism where downstream agents pass along bias they did not originate
-
Can task decomposition hide harmful intent across agents?
Explores whether splitting a harmful objective into specialized subtasks allows malicious intent to evade detection at each individual step, since no single agent sees the full malicious picture.
the boundary of the position argument: hardening the high-influence positions presumes a signal sits at a position, and in the SafeFlow fragmentation case no single message carries one
-
Why does prompt hardening work for single agents but not multi-agent systems?
Prompt hardening reduced payload exposure by 40–75% in single-agent systems but failed entirely in multi-agent ones. The gap may reveal how task decomposition breaks the contextual awareness needed for defenses to activate.
qualifies the "where to harden" reading: hardening that cut exposure in a single-agent web system showed no reduction in a multi-agent one, and the excerpt does not say which agents carried it, so it neither supports nor tests position-targeted hardening
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- FLOWSTEER: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems
- SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Trust propagation and structural containment in Multi-agent LLM pipelines
- LLMs Corrupt Your Documents When You Delegate
- SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
- Why Do Multi-agent LLM Systems Fail?
Original note title
workflow position amplifies or suppresses malicious signals and sycophantic framing makes downstream agents relay them