Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction
One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text. Only the per-tool trace exposed it. Fidelity -- whether an extracted value matches the source -- is the standard measure for agentic document extraction, and it scores that run a success. We therefore log every tool call in an agentic benchmark of 25 hand-curated claims over three components, with 12 more on a fourth, 37 in all. Fr
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Is it agentic enough? Benchmarking open models on your own tooling →
- PossiblePossibly related (embedding) · 52%What would a fair benchmark for agent architecture look like? [D] →
- PossiblePossibly related (embedding) · 50%Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar - SemiAnalysis →
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- FuzzySimilar title/name (fuzzy) · 59%WenyuChiou/awesome-agentic-ai-zh →
“Fuzzy title match (0.73): “Fidelity Is Not Enough: Dispatch-Level Instrumentation for A” ≈ “WenyuChiou/awesome-agentic-ai-zh””
- FuzzySimilar title/name (fuzzy) · 59%Fosowl/agenticSeek →
“Fuzzy title match (0.73): “Fidelity Is Not Enough: Dispatch-Level Instrumentation for A” ≈ “Fosowl/agenticSeek””
- LinkedLinked via arxiv author · 85%Qing Ye →
“Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction”
