Topic

Lua

50 items across the graph — tagged with Lua.

From the graph · 50

repo
langfuse/langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, Op…

repo
mlflow/mlflow

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-…

repo
promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl…

repo
comet-ml/opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d…

repo
Tencent/WeKnora

Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

repo
nmap/nmap

Nmap - the Network Mapper. Github mirror of official SVN repository.

repo
open-compass/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets…

repo
jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤…

repo
Helicone/helicone

🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

repo
Giskard-AI/giskard-oss

🐢 Open-Source Evaluation & Testing library for LLM Agents

repo
Marker-Inc-Korea/AutoRAG

AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.

repo
Kiln-AI/Kiln

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

repo
Tencent/AI-Infra-Guard

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

repo
open-compass/VLMEvalKit

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

repo
langwatch/langwatch

The platform for LLM evaluations and AI agent testing

repo
truera/trulens

Evaluation and Tracking for LLM Experiments and AI Agents

repo
modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

repo
onestardao/WFGY

WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Glo…

repo
trpc-group/trpc-agent-go

A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

repo
codota/tabnine-vscode

Visual Studio Code client for Tabnine. https://marketplace.visualstudio.com/items?itemName=TabNine.tabnine-vscode

repo
cvs-health/uqlm

UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

repo
NVIDIA-NeMo/Gym

Evaluate and improve models and agents using environments

repo
JudgmentLabs/judgeval

The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

repo
Sahir619/fable-method

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

repo
nolabs-ai/deepfabric

Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline

repo
MigoXLab/dingo

Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool

repo
Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI…

repo
chrisliu298/awesome-llm-unlearning

A resource repository for machine unlearning in large language models

repo
TIGER-AI-Lab/ClawBench

Open-source benchmark for browser AI agents on daily tasks.

repo
PacificAI/langtest

Deliver safe & effective language models

repo
run-llama/ParseBench

ParseBench - A Document Parsing Benchmark for AI Agents

repo
CodeSoul-co/Hypha

Harness-oriented agent system framework for production-grade LLM agent applications

repo
sb-ai-lab/RePlay

A Comprehensive Framework for Building End-to-End Recommendation Systems with State-of-the-Art Models

repo
SAILResearch/awesome-ai-leaderboard

A curated list of awesome leaderboard-oriented resources for AI domain

repo
lechmazur/nyt-connections

Benchmark that evaluates LLMs using 759 NYT Connections puzzles extended with extra trick words

repo
arklexai/arksim

Find your agents errors be fore your real users do

repo
alopatenko/LLMEvaluation

A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the ado…

repo
ServiceNow/eva

A New End-to-end Framework for Evaluating Voice Agents

repo
EuroEval/EuroEval

The robust European language model benchmark.

repo
strands-agents/evals

A comprehensive evaluation framework for AI agents and LLM applications.

repo
saccofrancesco/deepshot

AI-powered NBA game outcome predictor that uses advanced team stats and trend-based features to forecast winners and track model performance

repo
mahmoudrabie/agentic-ai

Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.

repo
OpenDCAI/One-Eval

Automated system for LLM evaluation via agents. Doc as below:

repo
hwfengcs/DM-Code-Agent

Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay.

repo
primitive-bench/primitive-bench

The marketplace for verifiable AI outcomes: state a task, get a fixed price upfront, and Primitive Bench routes across tools to complete it, refunded if the out…

repo
SantanderAI/autoguardrails

Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

repo
DrVrej/VJ-Base

An addon for Garry’s Mod that provides a collection of bases to assist in creating different types of addons.

repo
hidai25/eval-view

Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.

repo
taoAIGC/AICompare

one click to open multi AI sites | 一键打开多个 AI 站点,查看 AI 结果

repo
bridge-mind/bridgebench

The world's #1 vibe coding benchmark — models go head-to-head on real engineering tasks, judged blind, ranked by Elo.

Related topics