Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

PACE: A Proxy for Agentic Capability Evaluation

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benc

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Implements

Covers

Implements (incoming)

authored (incoming)

Covers (incoming)

Related across the graph

repoLazyAGI/LazyLLMrepoOpenDataBox/MemoryDatapersonLindia TjuatjarepoGiskard-AI/giskard-osspersonJiarui LiunewsI built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]repomodular/modularrepoProtocol-Lattice/go-agentrepowaveix/pkg-ffagentreporuvnet/agenticowrepogo-appsec/toolboxnewsBenchmarks: AntLing-3.0-flash a hybrid-reasoning MoE model built for production-scale agents.newsShow HN: Benchmark your eng team's AI agent maturity in 5 minutesnewsNew LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]repoaikssen/llm-benchmark-for-devrepoaffaan-m/ECCrepoSandAI-org/MagiCompilerrepomaseval/MASEvalrepoeunomia-bpf/bpf-benchmarkrepoprime-radiant-inc/gauntletrepogenlayerlabs/subzeroclawnewsI benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloadsrepolangwatch/langwatchrepoalmide/almiderepoharvard-cns/orlarepomirror29/inalpharepoTauricResearch/TradingAgentsrepous/crwnewsBuilding AX evals that actually workpersonVincent LopersonLintang Sutawikarepolitefuse/litefusereporatel-ai/ratelpersonYunze Xiaorepotwaldin/honerepoFosowl/agenticSeekrepobenchopt/benchoptrepoAndyyyy64/whichllmrepotoxy4ny/redteam-ai-benchmarkpersonJiayi Gengrepoalgorithmicsuperintelligence/optillmreponomograph/jigrepoagentjido/llm_dbnewsScarfBench: Benchmarking AI Agents for Enterprise Java Framework MigrationpersonXiang Yuereporun-llama/ParseBenchrepoAlanFokCo/agentscope-gorepoOskarsEzerins/llm-benchmarksrepoopensandbox-group/OpenSandboxrepolotus-data/lotusrepoOpenDCAI/One-EvalpersonGraham NeubigrepoMetabuilder-Labs/tokenjampersonAditya Bharat SoninewsGoogle Cloud tests AI agents with ambiguity-based benchmarks - IT Brief Australiarepobosun-ai/swiftiderepojuliangeymonat-jpg/mothragrepobytechefhq/bytechefpersonYueqi Songrepobrowser-use/browser-userepomodelscope/evalscoperepohe-yufeng/AgentProberepoOpenDataBox/Workspace-BenchrepoOpenHands/OpenHandspersonDaniel Leerepoagent-toolsnewsAtomicwork And New Measure Open Source AI Benchmark For ITSM - Open Source For You

Topics