newsReddit r/LocalLLaMATrust 52 · CommunityPublished 28d agoLive · 28d ago
I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.
TL;DR: The result looks TOO GOOD TO BE TRUE — that's exactly why I'm posting it. I ran Ox alpha on the entire SWE-bench Verified-Mini set with the official mini-swe-agent scaffold — the same agent used for the swebench.com Bash-Only leaderboard. Result: 48/50 = 96% resolved, judged locally with the official SWE-bench Docker harness. For scale: Claude Fable 5 scores 95% with Anthropic's own tuned agent scaffold, and under my identical bash-only scaffold the best
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
- PossiblePossibly related (embedding) · 55%Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead →
