Agent Benchmark
2 items across the graph — tagged with Agent Benchmark.
From the graph · 2
repo
hidai25/eval-view
→repoRegression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
js-lee-AI/awesome-llm-agent-papers
→A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey…
