Evaluation Framework
8 items across the graph — tagged with Evaluation Framework.
From the graph · 8
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl…
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
A New End-to-end Framework for Evaluating Voice Agents
The robust European language model benchmark.
[NeurIPS 2024] Evaluation harness for SWT-Bench, a benchmark for evaluating LLM repository-level test-generation
LLM and agent evaluation for Java & Kotlin. Runs in JUnit and CI. Spring AI, LangChain4j, Koog, Embabel, and any LLM client.
TuRTLe: A Unified Evaluation of LLMs for RTL Generation 🐢 (MLCAD 2025, ACM TODAES 2026)
🧪💥 Evaluation framework based on Vitest, the testing framework you familiar with, for agents, models, and more.
