InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries
Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially determine the legal outcome. We introduce InsufficiencyBench, the first legal benchmark targeting query-side insufficiency: whether a model recognizes when a query lacks legally material information, identifies what is missing, and refrains from premature conclusions. We formalize a taxonomy of eight canonical missing-element categories across three structural failure modes---switch, gating, and fatal prerequisite--- and
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%See what legal professionals say about the role of AI and law - Thomson Reuters Legal Solutions →
- PossiblePossibly related (embedding) · 47%The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix →
- PossiblePossibly related (embedding) · 45%Prompts as privilege – Courts grapple with questions over protections for lawyers' and experts' AI use - Reuters →
- FuzzyOverlapping authors or contributors · 62%usestrix/strix →
“Shared author/contributor keys: vincent”
- LinkedLinked via arxiv author · 85%Samuel J. Vincent →
“InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries”
- LinkedLinked via arxiv author · 85%Daniel Calloway →
“InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries”
- LinkedLinked via arxiv author · 85%Fangyi Yu →
“InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries”
- LinkedLinked via arxiv author · 85%Andrew M. Bean →
“InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries”
