newsReddit r/MachineLearningTrust 72 · CommunityPublished 1mo agoLive · 1mo ago
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage [R]
submitted by /u/julian88888888 [link] [comments]
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions →
- LinkedLinked via unknownTraceLab: Characterizing Coding Agent Workloads for LLM Serving →
- LinkedLinked via unknownAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
- PossiblePossibly related (embedding) · 47%Coding-agents can replicate scientific machine learning papers →
- PossiblePossibly related (embedding) · 46%TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution →
- PossiblePossibly related (embedding) · 47%tomasz-tomczyk/crit-web →
- PossiblePossibly related (embedding) · 55%RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications →
- PossiblePossibly related (embedding) · 49%Codesteward/codesteward →
Covers
Covers (incoming)
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperCoding-agents can replicate scientific machine learning paperspaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolutionrepotomasz-tomczyk/crit-webpaperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task SpecificationsrepoCodesteward/codestewardpaperSynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training
Related across the graph
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperTraceLab: Characterizing Coding Agent Workloads for LLM ServingrepoCodesteward/codestewardpaperSynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking TrainingpaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolutionrepotomasz-tomczyk/crit-webpaperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task SpecificationspaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionspaperCoding-agents can replicate scientific machine learning papers
