repoGitHubTrust 82 · PrimaryPublished 24d agoLive · 21d ago
minghinmatthewlam/openbench
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%Is it agentic enough? Benchmarking open models on your own tooling →
- PossiblePossibly related (embedding) · 46%I want people here to open your eyes and note how Google did not sign the anti- open-source-model coalition →
Covers
Covers (incoming)
Related across the graph
newsI built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]newsI want people here to open your eyes and note how Google did not sign the anti- open-source-model coalitionnewsIs it agentic enough? Benchmarking open models on your own tooling
