Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matters for interpreting long-context model behavior. Yet existing detectors miss these heads by construction: they reward heads whose attended token matches the generated token, a literal-copy criterion that captures where a head reads but not what it writes through its output-value (OV) circuit, the very mechanism that carries non-literal retrieval. We introduce Logit-Contribution
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownBreakthrough in long-context efficiency announced →
- LinkedLinked via unknownAttention →
- LinkedLinked via arxiv author · 85%Aryo Pradipta Gema →
“Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads”
- LinkedLinked via arxiv author · 85%Beatrice Alex →
“Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads”
- LinkedLinked via arxiv author · 85%Pasquale Minervini →
“Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads”
- PossiblePossibly related (embedding) · 55%wbopan/flashtrace →
- PossiblePossibly related (embedding) · 49%Contrastive Decoding Diffing (CDD): recovering verbatim finetuning data from logits alone, no weight access needed[R] →
