repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 23d ago
OskarsEzerins/llm-benchmarks
Popular LLM benchmarks for ruby code generation
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 53%Helix-7B →
- PossiblePossibly related (embedding) · 53%Evaluate a model properly →
- PossiblePossibly related (embedding) · 51%Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
- PossiblePossibly related (embedding) · 51%DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R] →
- PossiblePossibly related (embedding) · 48%Mistral's open-source Leanstral 1.5 aces formal math benchmarks and catches real bugs in code - the-decoder.com →
- PossiblePossibly related (embedding) · 45%GraphBU: MILP Instance Generation with Graph-Native Block Units →
- PossiblePossibly related (embedding) · 54%Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction →
Implements
Related to
Covers
Covers (incoming)
Implements (incoming)
Related across the graph
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?paperCan LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment ReproductionpaperPACE: A Proxy for Agentic Capability EvaluationpaperGraphBU: MILP Instance Generation with Graph-Native Block UnitsmodelHelix-7BnewsDeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]tutorialEvaluate a model properlynewsMistral's open-source Leanstral 1.5 aces formal math benchmarks and catches real bugs in code - the-decoder.com
