PACE: A Proxy for Agentic Capability Evaluation
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benc
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%agent-tools →
- PossiblePossibly related (embedding) · 62%ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration →
- PossiblePossibly related (embedding) · 58%opensandbox-group/OpenSandbox →
- PossiblePossibly related (embedding) · 58%bytechefhq/bytechef →
- PossiblePossibly related (embedding) · 58%run-llama/ParseBench →
- PossiblePossibly related (embedding) · 49%affaan-m/ECC →
- PossiblePossibly related (embedding) · 50%lotus-data/lotus →
- PossiblePossibly related (embedding) · 48%waveix/pkg-ffagent →
