repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 21d ago
benchopt/benchopt
A framework for reproducible, comparable benchmarks
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 53%Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
- PossiblePossibly related (embedding) · 53%Introducing GeneBench-Pro →
- PossiblePossibly related (embedding) · 52%TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution →
- PossiblePossibly related (embedding) · 52%DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R] →
Implements
Covers
Related across the graph
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperPACE: A Proxy for Agentic Capability EvaluationpaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionnewsDeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]newsIntroducing GeneBench-Pro
