Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 1m ago

Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability audit over a finite behavioral policy class: given policies H, observation support O, and estimand $τ$, we test whether O separates every pair with different $τ$. The audit requires zero model calls and resolves our diagnostic case: base-only observation collapses seven frozen deterministic policies into one equivalence class; full support yields seven classes and no

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 51%Evaluate a model properly
  • FuzzySimilar title/name (fuzzy) · 84%mudler/LocalAI

    Fuzzy title match (0.92): “Beyond Local Accuracy: A Protocol-Level Identifiability Audi” ≈ “mudler/LocalAI”

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: luo

  • LinkedLinked via arxiv author · 85%Junhao Luo

    Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

  • LinkedLinked via arxiv author · 85%Kaining Huang

    Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

  • LinkedLinked via arxiv author · 85%Ziqi Sha

    Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

  • LinkedLinked via arxiv author · 85%Wenxuan Tang

    Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

  • LinkedLinked via arxiv author · 85%Xinwei Deng

    Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

Explains

Implements (incoming)

authored (incoming)

Related across the graph

Topics