newsReddit r/MachineLearningTrust 52 · CommunityPublished 22d agoLive · 21d ago
I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]
Built an open-source AI coding agent that was 7%–75% cheaper than a cold "claude -p" run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: $6.83, 207 turns AutoDev Studio: ~$1.70 for the same bug The full benchmark (including cases where it loses) is in the README. So what's different? Most AI coding agents re-explore a repository from scratch on
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 65%Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software →
- PossiblePossibly related (embedding) · 63%minghinmatthewlam/openbench →
- PossiblePossibly related (embedding) · 61%Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
- PossiblePossibly related (embedding) · 58%RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications →
- PossiblePossibly related (embedding) · 57%PACE: A Proxy for Agentic Capability Evaluation →
Covers
paperGovern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Softwarerepominghinmatthewlam/openbenchpaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task SpecificationspaperPACE: A Proxy for Agentic Capability Evaluation
Related across the graph
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperPACE: A Proxy for Agentic Capability EvaluationpaperGovern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native SoftwarepaperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specificationsrepominghinmatthewlam/openbench
