Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · 3d ago

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built around a persistent, cross-interaction rule memory: when an attack succeeds, the framework abstract

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B

    Fuzzy title match (0.73): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “AgentCore-8B”

  • LinkedLinked via arxiv author · 85%Tongyan Hu

    A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

  • LinkedLinked via arxiv author · 85%Bryan Hooi

    A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

  • FuzzySimilar title/name (fuzzy) · 59%bojieli/ai-agent-book

    Fuzzy title match (0.73): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “bojieli/ai-agent-book”

  • FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent

    Fuzzy title match (0.94): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “SWE-agent/SWE-agent”

  • FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent

    Fuzzy title match (0.94): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “zhayujie/CowAgent”

  • FuzzySimilar title/name (fuzzy) · 66%open-multi-agent/open-multi-agent

    Fuzzy title match (0.78): “A Self-Evolving Multi-Agent Framework Defense against LLM Ja” ≈ “open-multi-agent/open-multi-agent”

Has model

Covers

authored (incoming)

Implements (incoming)

Related across the graph

Topics