Topic

Evals

17 items across the graph — tagged with Evals.

From the graph · 17

repo
mastra-ai/mastra

Mastra is the modern TypeScript framework for AI-powered applications and agents.

repo
Kiln-AI/Kiln

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

repo
truera/trulens

Evaluation and Tracking for LLM Experiments and AI Agents

repo
ombharatiya/ai-system-design-guide

AI system design guide for engineers building production AI systems and evals.

repo
future-agi/future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Gu…

repo
ShenSeanChen/waku-agent

Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all in code built to stay legible as it grows.

repo
Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI…

repo
hud-evals/hud-python

RL environments + evals for AI agents. Define once, train anything.

repo
zapier/AutomationBench

A benchmark for evaluating AI agents on realistic business workflows

repo
inferock/inferock-bench

Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.

repo
minghinmatthewlam/openbench

Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, sp…

repo
spences10/my-pi

Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry

repo
vstorm-co/awesome-pydantic-ai

An opinionated list of awesome Pydantic-AI frameworks, libraries, software and resources.

repo
dustalov/evalica

Evalica, your favourite evaluation toolkit

repo
METR/hawk

Run Inspect AI evals in the cloud

repo
skillberry-ai/cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

repo
vieval-dev/vieval

🧪💥 Evaluation framework based on Vitest, the testing framework you familiar with, for agents, models, and more.

Related topics