EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. Prior efforts, centered on benchmark distillation, reduce the number of evaluation tasks but leave the cost of executing each retained task untouched. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task. Our key insight is that an agent's final outcome
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “EarlyEval: Cheaper Agent Evaluation via Early Outcome Predic” ≈ “AgentCore-8B””
- PossiblePossibly related (embedding) · 62%Beyond benchmarks: The 5 pillars of AI evaluation systems →
- PossiblePossibly related (embedding) · 57%What would a fair benchmark for agent architecture look like? [D] →
- FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent →
“Fuzzy title match (0.94): “EarlyEval: Cheaper Agent Evaluation via Early Outcome Predic” ≈ “SWE-agent/SWE-agent””
- FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent →
“Fuzzy title match (0.94): “EarlyEval: Cheaper Agent Evaluation via Early Outcome Predic” ≈ “zhayujie/CowAgent””
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: wan”
- FuzzySimilar title/name (fuzzy) · 59%iflytek/astron-agent →
“Fuzzy title match (0.73): “EarlyEval: Cheaper Agent Evaluation via Early Outcome Predic” ≈ “iflytek/astron-agent””
