Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 28d agoLive · 28d ago

I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.

TL;DR: The result looks TOO GOOD TO BE TRUE — that's exactly why I'm posting it. I ran Ox alpha on the entire SWE-bench Verified-Mini set with the official mini-swe-agent scaffold — the same agent used for the swebench.com Bash-Only leaderboard. Result: 48/50 = 96% resolved, judged locally with the official SWE-bench Docker harness. For scale: Claude Fable 5 scores 95% with Anthropic's own tuned agent scaffold, and under my identical bash-only scaffold the best

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

Related across the graph