Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. We propose an action-class-oriented diagnostic framework that decomposes multi-turn failures into two orthogonal modes: action-class miscalibration and action-execution failure. The framework operates over a four-class action space (TOOL_CALL/A
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Is it agentic enough? Benchmarking open models on your own tooling →
- FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action →
“Fuzzy title match (0.92): “Calibration is the Bottleneck: An Action-Class Diagnostic of” ≈ “liguodongiot/llm-action””
- FuzzyOverlapping authors or contributors · 62%Kong/kong →
“Shared author/contributor keys: kong”
- FuzzyOverlapping authors or contributors · 62%rasbt/LLMs-from-scratch →
“Shared author/contributor keys: yin”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- LinkedLinked via arxiv author · 85%Kangjia Zhao →
“Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling”
- LinkedLinked via arxiv author · 85%Jiajun Liu →
“Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling”
- LinkedLinked via arxiv author · 85%Haozhan Shen →
“Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling”
