Llm Evaluation
20 items across the graph — tagged with Llm Evaluation.
From the graph · 20
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, Op…
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-…
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl…
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d…
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤…
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
🐢 Open-Source Evaluation & Testing library for LLM Agents
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
Evaluation and Tracking for LLM Experiments and AI Agents
UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI…
A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the ado…
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
LLM Evaluation for Phoenix Apps
A design-of-experiments platform for evaluating compound AI systems - find which technique drives quality, by how much, and whether the difference is real.
Combine the results from a panel of models into an enhanced response
Runtime security for AI agents. In-process, zero dependencies, Apache 2.0.
Detect prompt regressions before they reach production — per-category accuracy scoring, deterministic validation, and False Improvement detection. Pure Python,…
