Topic

Llm Evaluation

20 items across the graph — tagged with Llm Evaluation.

From the graph · 20

repo
langfuse/langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, Op…

repo
mlflow/mlflow

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-…

repo
promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl…

repo
comet-ml/opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d…

repo
jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤…

repo
Helicone/helicone

🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

repo
Giskard-AI/giskard-oss

🐢 Open-Source Evaluation & Testing library for LLM Agents

repo
Marker-Inc-Korea/AutoRAG

AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.

repo
Tencent/AI-Infra-Guard

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

repo
truera/trulens

Evaluation and Tracking for LLM Experiments and AI Agents

repo
cvs-health/uqlm

UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

repo
JudgmentLabs/judgeval

The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

repo
Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI…

repo
alopatenko/LLMEvaluation

A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the ado…

repo
skillberry-ai/cap-evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

repo
ccarvalho-eng/aludel

LLM Evaluation for Phoenix Apps

repo
fabian-lu/Cafe

A design-of-experiments platform for evaluating compound AI systems - find which technique drives quality, by how much, and whether the difference is real.

repo
shrdgn/openfusion

Combine the results from a panel of models into an enhanced response

repo
delphisecurity/xaidr

Runtime security for AI agents. In-process, zero dependencies, Apache 2.0.

repo
Emmimal/prompt-regression-suite

Detect prompt regressions before they reach production — per-category accuracy scoring, deterministic validation, and False Improvement detection. Pure Python,…

Related topics