A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built around a persistent, cross-interaction rule memory: when an attack succeeds, the framework abstract
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “AgentCore-8B””
- PossiblePossibly related (embedding) · 52%Beyond LLMs: Creating Real-World AI Agents with Lang Chain Deep Agents - HackerNoon →
- LinkedLinked via arxiv author · 85%Tongyan Hu →
“A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks”
- LinkedLinked via arxiv author · 85%Bryan Hooi →
“A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks”
- FuzzySimilar title/name (fuzzy) · 59%bojieli/ai-agent-book →
“Fuzzy title match (0.73): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “bojieli/ai-agent-book””
- FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent →
“Fuzzy title match (0.94): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “SWE-agent/SWE-agent””
- FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent →
“Fuzzy title match (0.94): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “zhayujie/CowAgent””
- FuzzySimilar title/name (fuzzy) · 66%open-multi-agent/open-multi-agent →
“Fuzzy title match (0.78): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “open-multi-agent/open-multi-agent””
