Read original ↗
paperarXivTrust 82 · PrimaryPublished 25d agoLive · 24d ago

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factual completeness, the missing half of factuality, is harder: it requires enumerating the full set of facts a complete answer should contain, and these facts rarely form a flat list. They often involve open-ended sets where coverage is what matters, ordered processes, and relationsh

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 47%New benchmark exposes reasoning gaps in top models
  • LinkedLinked via arxiv author · 85%Xilun Chen

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

  • LinkedLinked via arxiv author · 85%Zhaleh Feizollahi

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

  • LinkedLinked via arxiv author · 85%Ross Goodwin

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

  • LinkedLinked via arxiv author · 85%Seungwhan Moon

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

  • LinkedLinked via arxiv author · 85%Scott Yih

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

  • LinkedLinked via arxiv author · 85%Pinar Donmez

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

  • LinkedLinked via arxiv author · 85%Babak Damavandi

    Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Covers

authored (incoming)

Implements (incoming)

Related across the graph

Topics