repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 17h ago
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 66%LangChain Engineer Introduces Harbor for Complex AI Agent Evaluation - TechGig →
- PossiblePossibly related (embedding) · 64%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 61%AutoTrainess: Teaching Language Models to Improve Language Models Autonomously →
- PossiblePossibly related (embedding) · 58%SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions →
- PossiblePossibly related (embedding) · 57%AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents →
- PossiblePossibly related (embedding) · 54%I built a native Reddit app where a council of 5 AI agents debate and roast your project ideas →
- PossiblePossibly related (embedding) · 55%ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents →
- PossiblePossibly related (embedding) · 58%LLM-as-a-Verifier: A General-Purpose Verification Framework →
Covers
Implements
Covers (incoming)
newsI built a native Reddit app where a council of 5 AI agents debate and roast your project ideasnewsGoogle Cloud tests AI agents with ambiguity-based benchmarks - IT Brief AustralianewsTest solution uses agentic AI to turn natural language prompts into customized instruments - Military Embedded SystemsnewsHow to test agent skills without hitting real APIsnewsDevs shipping AI agents what does your security testing look like ?newsAI can’t simulate human preferences - new study tests LLMs against thousands of real usersnewsAgentic AI and cybersecurity, the story so farnewsEnterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them
Implements (incoming)
paperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperDoomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Related across the graph
paperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentsnewsLangChain Engineer Introduces Harbor for Complex AI Agent Evaluation - TechGignewsAI can’t simulate human preferences - new study tests LLMs against thousands of real userspaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperPACE: A Proxy for Agentic Capability EvaluationpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentsnewsEnterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify themnewsDevs shipping AI agents what does your security testing look like ?paperDoomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe CascadenewsGoogle Cloud tests AI agents with ambiguity-based benchmarks - IT Brief AustraliapaperAutoTrainess: Teaching Language Models to Improve Language Models AutonomouslynewsTest solution uses agentic AI to turn natural language prompts into customized instruments - Military Embedded SystemspaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsnewsAgentic AI and cybersecurity, the story so farnewsI built a native Reddit app where a council of 5 AI agents debate and roast your project ideasnewsHow to test agent skills without hitting real APIs
