Evaluation
50 items across the graph — tagged with Evaluation.
From the graph · 50
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, Op…
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-…
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl…
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d…
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets…
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤…
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
🐢 Open-Source Evaluation & Testing library for LLM Agents
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
The platform for LLM evaluations and AI agent testing
Evaluation and Tracking for LLM Experiments and AI Agents
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Glo…
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection
Evaluate and improve models and agents using environments
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI…
A resource repository for machine unlearning in large language models
Open-source benchmark for browser AI agents on daily tasks.
Deliver safe & effective language models
ParseBench - A Document Parsing Benchmark for AI Agents
A Comprehensive Framework for Building End-to-End Recommendation Systems with State-of-the-Art Models
A curated list of awesome leaderboard-oriented resources for AI domain
Benchmark that evaluates LLMs using 759 NYT Connections puzzles extended with extra trick words
Find your agents errors be fore your real users do
A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the ado…
The robust European language model benchmark.
A New End-to-end Framework for Evaluating Voice Agents
A comprehensive evaluation framework for AI agents and LLM applications.
AI-powered NBA game outcome predictor that uses advanced team stats and trend-based features to forecast winners and track model performance
Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.
Automated system for LLM evaluation via agents. Doc as below:
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay.
The marketplace for verifiable AI outcomes: state a task, get a fixed price upfront, and Primitive Bench routes across tools to complete it, refunded if the out…
Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
one click to open multi AI sites | 一键打开多个 AI 站点,查看 AI 结果
The world's #1 vibe coding benchmark — models go head-to-head on real engineering tasks, judged blind, ranked by Elo.
RouterArena: An open framework for evaluating LLM routers with standardized datasets, metrics, an automated framework, and a live leaderboard.
Make AI work for Everyone - Monitoring and governing for your AI/ML
[NeurIPS 2024] Evaluation harness for SWT-Bench, a benchmark for evaluating LLM repository-level test-generation
Litefuse - Agent Observability and Evaluation Platform
