repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
OpenDCAI/One-Eval
Automated system for LLM evaluation via agents. Doc as below:
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 53%Evaluate a model properly →
- PossiblePossibly related (embedding) · 53%Hypothetically speaking... →
- PossiblePossibly related (embedding) · 52%AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents →
- PossiblePossibly related (embedding) · 52%A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis →
- PossiblePossibly related (embedding) · 59%LLM-as-a-Verifier: A General-Purpose Verification Framework →
- PossiblePossibly related (embedding) · 47%LLM Judges Can Be Too Generous When There Is No Reference Answer →
- PossiblePossibly related (embedding) · 48%Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026 →
Implements
Related to
Covers
Implements (incoming)
Covers (incoming)
Related across the graph
newsCan a MUD evaluate LLMs? A $99 proof of conceptpaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperLLM Judges Can Be Too Generous When There Is No Reference AnswerpaperPACE: A Proxy for Agentic Capability EvaluationnewsEvals are the new PRD, Expedia’s AI chief tells VB Transform 2026newsHypothetically speaking...paperA Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure DiagnosistutorialEvaluate a model properly
