The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- PossiblePossibly related (embedding) · 55%thu-pacman/chitu →
- LinkedLinked via arxiv author · 85%Yanzhe Zhang →
“The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context”
- LinkedLinked via arxiv author · 85%Sanmi Koyejo →
“The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context”
- LinkedLinked via arxiv author · 85%Diyi Yang →
“The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context”
