Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that joi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%What would a fair benchmark for agent architecture look like? [D] →
- PossiblePossibly related (embedding) · 54%New Open Source Benchmark Scores AI Agents on Their Ability to Learn and Perform Complex Actions - WBOC TV →
- LinkedLinked via arxiv author · 85%Xinshuai Guo →
“Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking”
- LinkedLinked via arxiv author · 85%Junjie Wu →
“Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking”
- LinkedLinked via arxiv author · 85%Dolly Deng →
“Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking”
- LinkedLinked via arxiv author · 85%Yinghui Li →
“Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking”
- LinkedLinked via arxiv author · 85%Hai-Tao Zheng →
“Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking”
- LinkedLinked via arxiv author · 85%Suncong Zheng →
“Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking”
