newsGoogle News — LLMTrust 62 · AggregatorPublished 1mo agoLive · 1mo ago
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models - Apple Machine Learning Research
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Apple Machine Learning Research
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models →
- PossiblePossibly related (embedding) · 52%Harnessing Textual Refusal Directions for Multimodal Safety →
- PossiblePossibly related (embedding) · 49%Faithfulness to Refusal: A Causal Audit of Neuron Selectors →
- PossiblePossibly related (embedding) · 48%HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety →
- PossiblePossibly related (embedding) · 49%Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks →
- PossiblePossibly related (embedding) · 48%Neural Collapse Is Forbidden: Information Floors in Language Models →
Covers
paperRobust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language ModelspaperHarnessing Textual Refusal Directions for Multimodal SafetypaperFaithfulness to Refusal: A Causal Audit of Neuron SelectorspaperHaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety
Covers (incoming)
Related across the graph
paperNeural Collapse Is Forbidden: Information Floors in Language ModelspaperHarnessing Textual Refusal Directions for Multimodal SafetypaperFaithfulness to Refusal: A Causal Audit of Neuron SelectorspaperRobust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language ModelspaperContravariance Theory: Strong Alignment for Minimal Solutions to Hard TaskspaperHaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety
