Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields t
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Adam Karvonen →
“Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments”
- LinkedLinked via arxiv author · 85%Euan Ong →
“Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments”
- LinkedLinked via arxiv author · 85%Subhash Kantamneni →
“Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments”
- LinkedLinked via arxiv author · 85%Samuel Marks →
“Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments”
