repoGitHubTrust 82 · PrimaryPublished 3mo agoLive · 2mo ago
eval-harness-plus
An extensible evaluation harness for LLMs.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownEvaluate a model properly →
- LinkedLinked via unknownEvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures →
- PossiblePossibly related (embedding) · 45%Online Safety Monitoring for LLMs →
- PossiblePossibly related (embedding) · 49%When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability →
- PossiblePossibly related (embedding) · 50%Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026 →
Related to
Covers (incoming)
Implements (incoming)
Related across the graph
newsHow're you deploying LLMs in production now-a-days? What's the best and most affordable way? [D]paperEvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety FailuresnewsEvals are the new PRD, Expedia’s AI chief tells VB Transform 2026paperOnline Safety Monitoring for LLMstutorialEvaluate a model properlypaperWhen the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
