Faithfulness to Refusal: A Causal Audit of Neuron Selectors
Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly. We address this with two paired audits built on one-shot neuron-row zeroing. We first audit selectors at the language-modeling level: attribution methods substantially outperform activation and magnitude-based baselines at identifying dispensable rows across five LLMs. We then adapt the same intervention into a behavior test by driving it with a contrastive h
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Ananth Eswar →
“Faithfulness to Refusal: A Causal Audit of Neuron Selectors”
- LinkedLinked via arxiv author · 85%Pratinav Seth →
“Faithfulness to Refusal: A Causal Audit of Neuron Selectors”
- LinkedLinked via arxiv author · 85%Utsav Avaiya →
“Faithfulness to Refusal: A Causal Audit of Neuron Selectors”
- LinkedLinked via arxiv author · 85%Vinay Kumar Sankarapu →
“Faithfulness to Refusal: A Causal Audit of Neuron Selectors”
- PossiblePossibly related (embedding) · 49%A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models - Apple Machine Learning Research →
