Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission. We audit these issues across the three benchmarks. First, we replay the official reference patches for 740 code opti

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

authored (incoming)

Implements (incoming)

Related across the graph

personYuling ShipersonZhensu SunnewsI built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]repoAgent-Field/pr-afrepolemoncrowhq/lemoncrownewsOrnith-1.0: self-improving open-source models for agentic codingnewsDoes code cleanliness affect coding agents? A controlled minimal-pair studyrepoaffaan-m/ECCpersonLingxiao JiangnewsI benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloadspersonZhi ChennewsForrester Consulting: GitLab Duo Agent Platform delivers 400% ROIrepobenchopt/benchoptnewsBetter tools made Copilot code review worse. Here’s how we actually improved it.newsGoogle's Agentic Peer-Reviewer Handled ~10K Papers at ICML/STOC — Formal Research Paper Now Out [R]newsSenior SWE-Bench: open-source benchmark that assesses agents as senior engineersnewsScarfBench: Benchmarking AI Agents for Enterprise Java Framework MigrationrepoUnity-Technologies/ml-agentsrepoNirDiamant/GenAI_AgentsrepoOskarsEzerins/llm-benchmarkspersonDavid LonewsDeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]repogolobokov.misha/llm-review-agentsnewsA fully local, self-hosted repo index for coding agents (Rust, MIT, runs offline)repogoogle-research/google-researchrepoBerriAI/litellmnewsREAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage [R]newsMistral's open-source Leanstral 1.5 aces formal math benchmarks and catches real bugs in code - the-decoder.com

Topics