Benchmarks
22 items across the graph · 3 news stories — tagged with Benchmarks.
Latest news
GLM 5.2 beats Claude in our benchmarks
Article URL: https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/ Comments URL: https://news.ycombinator.com/item?id=48709670 Points: 1011 # Comments: 466
Read full story →More news · 2
Semgrep: GLM 5.2 beats Claude in our Cyber Benchmarks
Article URL: https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/ Comments URL: https://news.ycombinator.com/item?id=48709670 Points: 56 # Comments: 20
Read full story →New benchmark exposes reasoning gaps in top models
A harder evaluation suite shows even leading models struggle on multi-hop tasks.
Read full story →From the graph · 19
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it ins…
Evaluate and improve models and agents using environments
Mother of All BCI Benchmarks
Deliver safe & effective language models
TUI and CLI for browsing AI models, benchmarks, coding agents, and statuses for AI providers.
Minecraft-style voxel benchmark for comparing AI models (Arena + Sandbox)
Curated Claude Fable 5 use cases, tutorials, integrations, demos, and benchmark evidence with source links. Access Claude Fable 5 exclusively via MuAPI.
Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.
Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
A benchmark for evaluating AI agents on realistic business workflows
scikit-learn_bench benchmarks various implementations of machine learning algorithms across data analytics frameworks. It currently support the scikit-learn, DA…
[NeurIPS'25 | CVPR'26] The official repo of OralGPT & MMOral Bench.
Popular LLM benchmarks for ruby code generation
Strands-based agents and harnesses for agentic benchmarks.
Experiments with reasoning models, training techniques, papers
All-in-one open-source framework & middleware for enterprise-grade multi-tenant and multi-tier XaaS Services development
Avoid common pitfalls when benchmarking LLMs.
A dashboard for tracking model evaluations over time.
An extensible evaluation harness for LLMs.
