When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself: split fragments can still leak suspicious tokens or provenance edges. The hard case is \emph{local benignness}. No fragment carries the harm, and what is left looks like or
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%A system-level approach to prompt injection: separating instruction and data channels in LLM agents [P] →
- PossiblePossibly related (embedding) · 51%SentryCode: Real-time Auditor + Honeytokens for AI Coding Agents [P] →
- PossiblePossibly related (embedding) · 50%Vishisht16/Humane-Proxy →
- PossiblePossibly related (embedding) · 48%Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign) →
- PossiblePossibly related (embedding) · 48%huhusmang/Awesome-LLMs-for-Vulnerability-Detection →
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “When Local Monitors Miss Compositional Harm: Diagnosing Dist” ≈ “AgentCore-8B””
- LinkedLinked via arxiv author · 85%Yibo Hu →
“When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems”
- LinkedLinked via arxiv author · 85%Ren Wang →
“When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems”
