CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose \textbf{C}ontinuous \textbf{L}at\textbf{E}nt \textbf{A}dapter \textbf{R}outing (CLEAR), a conditional safety adaptation framework that uses a lightweight hidden-state gate to continuously control the activation strength of a safety low-rank adapter. CLEAR aims to reduce harmful completions while avoiding unnecessary changes to the frozen backbone that could degrade performance on benign prompts. Ex
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 51%What does "Safe AI" look like? [D] →
- PossiblePossibly related (embedding) · 49%Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign) →
- PossiblePossibly related (embedding) · 49%A system-level approach to prompt injection: separating instruction and data channels in LLM agents [P] →
- PossiblePossibly related (embedding) · 49%Safety and alignment in an era of long-horizon models →
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%BerriAI/litellm →
“Shared author/contributor keys: jiang”
