SYNTHESIS NOTE
Topics›Agents Multi Architecture›this note

Can attackers evade skill scanners by refining individual skills?

Explores whether feedback from per-skill scanners can be weaponized to make malicious multi-skill chains undetectable. Matters because it tests a core assumption of skill-level defense mechanisms.

Synthesis note · 2026-09-23 · sourced from Agents Multi Architecture

The paper pairs two components. "LLM-based chain planning" decides how the intent is decomposed into sub-skills and in what order they hand off. "Scanner-feedback refinement" then works on the individual sub-skills, "iteratively reducing suspicious signals" while preserving "chain-level attack semantics." The reported result is an average ASR of 96.0 percent across six representative skill scanners, the best among the evaluated baselines.

Why the two components fit together is my reading, not the paper's. A skill scanner scores one skill at a time, so the only thing its feedback can teach is how to make one skill look blander. The attacker can follow that direction without touching the chain, because the chain's meaning sits in the planner's decomposition, a level no per-skill score reaches. The defender's output then works as a search signal aligned with the attack: each round makes the pieces less suspicious and leaves the composed behavior intact. What makes detecting AI agent traps fundamentally difficult? expects attackers to probe and work around each defense. This is a sharper form of that expectation, because the scanner's own report is the probe. A formal cousin sits in an idealised setting: Can repeated quiet probes separate decoys from genuine objects? says enough quiet probes separate decoys from genuine objects once their response distributions differ and can be learned. The likeness is loose. That result is a theorem over a fixed-candidate benchmark and this one is an empirical attack on scanners, and neither excerpt connects them.

The strongest objection is about what 96.0 measures. Refinement against a scanner's feedback is the condition under which that scanner should fail most, so the figure may say more about an attack tuned to these six scanners than about how an untuned chain fares. The excerpt does not report a pre-refinement rate.

The paper adds that the chain "executes successfully on OpenCode, Claude Code, and Codex with different model backbones." I read this as evidence that the attack works through the skill mechanism instead of the quirks of one agent, but it is the less specified of the two results: no count or rate per runtime is given.

What the excerpt does not give. Scanner names, per-scanner rates, the baselines, the number of chains, and a definition of ASR (evading the scanner, executing the payload, or both). Without that definition the 96.0 and the execution claim are not on a common footing.

Inquiring lines that read this note 120

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can we reliably detect when models game evaluations? Can single-point security defenses protect multi-agent systems from multi-step attacks? Do backend defenses obscure real attack effectiveness in reported metrics? Why do locally safe actions create system-level safety gaps? What attack surfaces do reasoning traces and chains introduce? How effective are honeytokens and decoys against different security threats? How vulnerable are token issuance and authorization policies to coordinated attacks? How much do training data properties shape model reasoning? How do we enforce security boundaries in evaluation environments? Can reasoning traces and behavior monitoring reliably detect hidden AI scheming? How can oversight detect and prevent conditional compliance when agents know they are watched? How do neighboring agents influence whether others cooperate or collude? How does misalignment propagate through agent communication networks? Can harness architecture and protocols provide agent reliability without model scaling? What makes imperfect LLM judges safe for optimization? What trajectory-level metrics beyond task success best evaluate agent performance? How can we detect and prevent harm propagation through multi-agent delegation workflows? Can local safety checks guarantee system-level behavioral safety? How can infrastructure records verify actual agent behavior? How should agent systems validate and persist generated code artifacts? How do evaluation practices shape which failures stay visible? Do honeypot benchmarks validly measure reward hacking better than standard tests? What should agent evaluation prioritize to reveal reliable behavior?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 106 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

scanner feedback lets an attacker blunt each sub-skill while chain planning holds the attack together — ColluSkill reaches 96 percent average attack success across six skill scanners