Topic

Benchmarks

22 items across the graph · 3 news stories — tagged with Benchmarks.

Latest news

NewsHacker NewsLive · 1mo ago

GLM 5.2 beats Claude in our benchmarks

Article URL: https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/ Comments URL: https://news.ycombinator.com/item?id=48709670 Points: 1011 # Comments: 466

Read full story →

More news · 2

From the graph · 19

repo
Andyyyy64/whichllm

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it ins…

repo
NVIDIA-NeMo/Gym

Evaluate and improve models and agents using environments

repo
NeuroTechX/moabb

Mother of All BCI Benchmarks

repo
PacificAI/langtest

Deliver safe & effective language models

repo
reyamira/models

TUI and CLI for browsing AI models, benchmarks, coding agents, and statuses for AI providers.

repo
Ammaar-Alam/minebench

Minecraft-style voxel benchmark for comparing AI models (Arena + Sandbox)

repo
Anil-matcha/awesome-claude-fable-5

Curated Claude Fable 5 use cases, tutorials, integrations, demos, and benchmark evidence with source links. Access Claude Fable 5 exclusively via MuAPI.

repo
mahmoudrabie/agentic-ai

Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.

repo
thenicolas1894/awesome-claude-fable-5-prompt-vault

Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks

repo
zapier/AutomationBench

A benchmark for evaluating AI agents on realistic business workflows

repo
IntelPython/scikit-learn_bench

scikit-learn_bench benchmarks various implementations of machine learning algorithms across data analytics frameworks. It currently support the scikit-learn, DA…

repo
isjinghao/OralGPT

[NeurIPS'25 | CVPR'26] The official repo of OralGPT & MMOral Bench.

repo
OskarsEzerins/llm-benchmarks

Popular LLM benchmarks for ruby code generation

repo
strands-labs/benchmark-harnesses

Strands-based agents and harnesses for agentic benchmarks.

repo
benjaminzwhite/reasoning-models

Experiments with reasoning models, training techniques, papers

repo
constructorfabric/gears-rust

All-in-one open-source framework & middleware for enterprise-grade multi-tenant and multi-tier XaaS Services development

tutorial
Evaluate a model properly

Avoid common pitfalls when benchmarking LLMs.

tool
EvalBoard

A dashboard for tracking model evaluations over time.

repo
eval-harness-plus

An extensible evaluation harness for LLMs.

Related topics