repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%LangChain Engineer Introduces Harbor for Complex AI Agent Evaluation - TechGig →
- PossiblePossibly related (embedding) · 60%Predicting model behavior before release by simulating deployment →
- PossiblePossibly related (embedding) · 54%Helping build shared standards for advanced AI →
- PossiblePossibly related (embedding) · 50%GLM 5.2: New Chinese AI Model Nearly Matches Anthropic and OpenAI in Benchmarks - News and Statistics - IndexBox →
- PossiblePossibly related (embedding) · 55%Mistral Open-Sources AI Model That Can Verify Code and Mathematical Proofs - ProPakistani →
- PossiblePossibly related (embedding) · 49%How to test agent skills without hitting real APIs →
- PossiblePossibly related (embedding) · 49%How to test agent experience changes without shipping them →
Covers
Covers (incoming)
newsGLM 5.2: New Chinese AI Model Nearly Matches Anthropic and OpenAI in Benchmarks - News and Statistics - IndexBoxnewsMistral Open-Sources AI Model That Can Verify Code and Mathematical Proofs - ProPakistaninewsHow to test agent skills without hitting real APIsnewsHow to test agent experience changes without shipping them
Related across the graph
newsMistral Open-Sources AI Model That Can Verify Code and Mathematical Proofs - ProPakistaninewsLangChain Engineer Introduces Harbor for Complex AI Agent Evaluation - TechGignewsGLM 5.2: New Chinese AI Model Nearly Matches Anthropic and OpenAI in Benchmarks - News and Statistics - IndexBoxnewsPredicting model behavior before release by simulating deploymentnewsHow to test agent experience changes without shipping themnewsHelping build shared standards for advanced AInewsHow to test agent skills without hitting real APIs
