Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit ana
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Attention →
- PossiblePossibly related (embedding) · 45%Language Models Can Control Their Own Attention [R] →
- FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference →
“Fuzzy title match (0.92): “Influence Score and Transformers interpretability: Measure o” ≈ “xorbitsai/inference””
- FuzzySimilar title/name (fuzzy) · 84%huggingface/transformers →
“Fuzzy title match (0.92): “Influence Score and Transformers interpretability: Measure o” ≈ “huggingface/transformers””
- LinkedLinked via arxiv author · 85%Lisa Bouger →
“Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time”
- LinkedLinked via arxiv author · 85%Yannick Teglia →
“Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time”
- LinkedLinked via arxiv author · 85%Philippe Loubet Moundi →
“Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time”
