Read original ↗
paperarXivTrust 82 · PrimaryPublished yesterdayLive · 1h ago

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success using simple and direct factual recall. This framing fails to capture a key requirement of unlearning, namely the ability to eliminate harmful behaviors while preserving benign and beneficial knowledge. We argue that effective unlearning must operate at the level of concepts, ensuring c

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Sahil Kale

    ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

  • LinkedLinked via arxiv author · 85%Ian Harris

    ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

authored (incoming)

Related across the graph

Topics