Read original ↗
paperarXivTrust 82 · PrimaryPublished 7d agoLive · 2d ago

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit ana

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 51%Attention
  • PossiblePossibly related (embedding) · 45%Language Models Can Control Their Own Attention [R]
  • FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference

    Fuzzy title match (0.92): “Influence Score and Transformers interpretability: Measure o” ≈ “xorbitsai/inference”

  • FuzzySimilar title/name (fuzzy) · 84%huggingface/transformers

    Fuzzy title match (0.92): “Influence Score and Transformers interpretability: Measure o” ≈ “huggingface/transformers”

  • LinkedLinked via arxiv author · 85%Lisa Bouger

    Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

  • LinkedLinked via arxiv author · 85%Yannick Teglia

    Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

  • LinkedLinked via arxiv author · 85%Philippe Loubet Moundi

    Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

Related to

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics