repoGitHubTrust 82 · PrimaryPublished yesterdayLive · 14h ago
JudgmentLabs/judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Evaluating AI Agents: A production blueprint with Strands and AgentCore →
- PossiblePossibly related (embedding) · 55%The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway →
