Evals
17 items across the graph — tagged with Evals.
From the graph · 17
Mastra is the modern TypeScript framework for AI-powered applications and agents.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Evaluation and Tracking for LLM Experiments and AI Agents
AI system design guide for engineers building production AI systems and evals.
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Gu…
Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all in code built to stay legible as it grows.
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI…
RL environments + evals for AI agents. Define once, train anything.
A benchmark for evaluating AI agents on realistic business workflows
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, sp…
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
An opinionated list of awesome Pydantic-AI frameworks, libraries, software and resources.
Evalica, your favourite evaluation toolkit
Run Inspect AI evals in the cloud
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
🧪💥 Evaluation framework based on Vitest, the testing framework you familiar with, for agents, models, and more.
