Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 19h ago

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Adam Karvonen

    Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

  • LinkedLinked via arxiv author · 85%Euan Ong

    Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

  • LinkedLinked via arxiv author · 85%Subhash Kantamneni

    Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

  • LinkedLinked via arxiv author · 85%Samuel Marks

    Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

authored (incoming)

Related across the graph

Topics