Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission. We audit these issues across the three benchmarks. First, we replay the official reference patches for 740 code opti
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownREAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage [R] →
- LinkedLinked via unknownScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration →
- LinkedLinked via unknownDeepSWE: new benchmark looking at how well today's frontier models can actually write code [R] →
- LinkedLinked via unknownOrnith-1.0: self-improving open-source models for agentic coding →
- LinkedLinked via unknownSenior SWE-Bench: open-source benchmark that assesses agents as senior engineers →
- LinkedLinked via arxiv author · 85%Zhi Chen →
“Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?”
- LinkedLinked via arxiv author · 85%Zhensu Sun →
“Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?”
