What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably imp
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Evaluate a model properly →
- PossiblePossibly related (embedding) · 51%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 48%Benchmarking Self-Hosted Gemma 2 9B vs. Frontier APIs: The FP8 Quantization Prefill Tax and VRAM Realities on an NVIDIA L4 [P] →
- PossiblePossibly related (embedding) · 47%A 35%-accurate model that still ranked well, and adding more features made it worse (point-in-time equity backtest) [P] →
- PossiblePossibly related (embedding) · 45%We compared different LLMs on IMO 2026 [R] →
- LinkedLinked via arxiv author · 85%Xiaonan Xu →
“What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations”
- LinkedLinked via arxiv author · 85%Wenjing Wu →
“What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations”
