repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 14d ago
sierra-research/tau2-bench
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use →
- PossiblePossibly related (embedding) · 62%UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks →
- PossiblePossibly related (embedding) · 62%ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents →
- PossiblePossibly related (embedding) · 61%SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions →
- PossiblePossibly related (embedding) · 59%PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents →
- PossiblePossibly related (embedding) · 60%MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents →
- PossiblePossibly related (embedding) · 54%Most AI agents have no concept of opportunity cost [D] →
Implements
paperCan Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool UsepaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Implements (incoming)
Covers (incoming)
Related across the graph
paperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentsnewsMost AI agents have no concept of opportunity cost [D]paperMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentspaperCan Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool UsepaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
