Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence b
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “Efficient SWE Agent Benchmarking via Trajectory-Aware Evalua” ≈ “AgentCore-8B””
- PossiblePossibly related (embedding) · 59%Beyond benchmarks: The 5 pillars of AI evaluation systems →
- PossiblePossibly related (embedding) · 57%What would a fair benchmark for agent architecture look like? [D] →
- PossiblePossibly related (embedding) · 56%Accelerating software delivery with agentic QA automation using Amazon Nova Act – Part 2 →
- FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent →
“Fuzzy title match (0.94): “Efficient SWE Agent Benchmarking via Trajectory-Aware Evalua” ≈ “SWE-agent/SWE-agent””
- FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent →
“Fuzzy title match (0.94): “Efficient SWE Agent Benchmarking via Trajectory-Aware Evalua” ≈ “zhayujie/CowAgent””
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
