Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a factor would change the output, and sufficiency, meaning that retaining it while removing other changeable information would preserve the output. We evalu
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%When an AI agent says “done” how do you know it actually happened? [P] →
- LinkedLinked via arxiv author · 85%Urja Pawar →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
- LinkedLinked via arxiv author · 85%Rajitha Ramanayake →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
- LinkedLinked via arxiv author · 85%Nabeel Kemal →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
- LinkedLinked via arxiv author · 85%Ashwin Kandath →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
- LinkedLinked via arxiv author · 85%Owen O'Neill →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
- LinkedLinked via arxiv author · 85%Guillaume Bourgeon →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
- LinkedLinked via arxiv author · 85%Houssem Chatbri →
“Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence”
