repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
zapier/AutomationBench
A benchmark for evaluating AI agents on realistic business workflows
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration →
- PossiblePossibly related (embedding) · 57%Advancing Efficient AI Workflows - Harvard University →
- PossiblePossibly related (embedding) · 58%PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents →
- PossiblePossibly related (embedding) · 61%Evolving from legacy BI to agentic AI at Tradeshift with Amazon Quick →
- PossiblePossibly related (embedding) · 53%Build specialized agent workflows for your business with Amazon Quick and NVIDIA NeMo Agent Toolkit →
- PossiblePossibly related (embedding) · 49%Where should deterministic business logic live in an AI-powered Azure architecture? →
- PossiblePossibly related (embedding) · 56%Agentic AI costs set to balloon fivefold by 2028 →
- PossiblePossibly related (embedding) · 62%How we built an AI agent for field associates with Red Hat AI →
Covers
Implements
Covers (incoming)
newsEvolving from legacy BI to agentic AI at Tradeshift with Amazon QuicknewsBuild specialized agent workflows for your business with Amazon Quick and NVIDIA NeMo Agent ToolkitnewsWhere should deterministic business logic live in an AI-powered Azure architecture?newsAgentic AI costs set to balloon fivefold by 2028newsHow we built an AI agent for field associates with Red Hat AInewsHow AWS Finance teams reclaimed hundreds of hours with Amazon QuicknewsShow HN: Benchmark your eng team's AI agent maturity in 5 minutesnewsAccelerating software delivery with agentic QA automation using Amazon Nova Act – Part 2newsScaling UX testing with Amazon Nova Act: A new approach to user flow analysisnewsThe real bottleneck for AI agents may be proving who they arenewsEnterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify themnewsScaling agentic workflows with native case management in Amazon Quick AutomatenewsBeyond benchmarks: The 5 pillars of AI evaluation systemsnewsThe agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
Implements (incoming)
Related across the graph
newsBuild specialized agent workflows for your business with Amazon Quick and NVIDIA NeMo Agent ToolkitpaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentsnewsThe real bottleneck for AI agents may be proving who they arenewsScaling agentic workflows with native case management in Amazon Quick AutomatenewsWhere should deterministic business logic live in an AI-powered Azure architecture?paperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TaskspaperAn Experimental Design Approach to Evaluating Agentic AI's Autonomous Model DiscoverynewsShow HN: Benchmark your eng team's AI agent maturity in 5 minutesnewsHow AWS Finance teams reclaimed hundreds of hours with Amazon QuickpaperDo AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionnewsEvolving from legacy BI to agentic AI at Tradeshift with Amazon QuicknewsEnterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify themnewsScaling UX testing with Amazon Nova Act: A new approach to user flow analysisnewsAdvancing Efficient AI Workflows - Harvard UniversitynewsScarfBench: Benchmarking AI Agents for Enterprise Java Framework MigrationnewsHow we built an AI agent for field associates with Red Hat AInewsAgentic AI costs set to balloon fivefold by 2028newsThe agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anywaynewsBeyond benchmarks: The 5 pillars of AI evaluation systemsnewsAccelerating software delivery with agentic QA automation using Amazon Nova Act – Part 2
