Read original ↗
paperarXivTrust 82 · PrimaryPublished 10d agoLive · 9d ago

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. We propose an action-class-oriented diagnostic framework that decomposes multi-turn failures into two orthogonal modes: action-class miscalibration and action-execution failure. The framework operates over a four-class action space (TOOL_CALL/A

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action

    Fuzzy title match (0.92): “Calibration is the Bottleneck: An Action-Class Diagnostic of” ≈ “liguodongiot/llm-action”

  • FuzzyOverlapping authors or contributors · 62%Kong/kong

    Shared author/contributor keys: kong

  • FuzzyOverlapping authors or contributors · 62%rasbt/LLMs-from-scratch

    Shared author/contributor keys: yin

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • LinkedLinked via arxiv author · 85%Kangjia Zhao

    Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

  • LinkedLinked via arxiv author · 85%Jiajun Liu

    Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

  • LinkedLinked via arxiv author · 85%Haozhan Shen

    Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Covers

Implements (incoming)

authored (incoming)

Covers (incoming)

Related across the graph

Topics