Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels. We use WIFA as a common data layer for two complementary fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and Anchored Group-Consistent
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 84%roboflow/supervision →
“Fuzzy title match (0.92): “Refusing Intent, Not Form: Wrapper-Based Intent-Group Superv” ≈ “roboflow/supervision””
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: luo”
- LinkedLinked via arxiv author · 85%Ping Wu →
“Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety”
- LinkedLinked via arxiv author · 85%Haibo Tong →
“Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety”
- LinkedLinked via arxiv author · 85%Feifei Zhao →
“Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety”
- LinkedLinked via arxiv author · 85%Shuhan Shen →
“Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety”
- LinkedLinked via arxiv author · 85%Yu Shi →
“Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety”
- LinkedLinked via arxiv author · 85%Yilin Zhao →
“Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety”
