Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · yesterday

TIGER-AI-Lab/ClawBench

Open-source benchmark for browser AI agents on daily tasks.

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

Implements (incoming)

Related across the graph

newsAI Innovators Adopt NVIDIA Vera — Why Max Single-Threaded CPU at Scale MatterspaperUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TasksnewsAI Agents Consume Up to 136 Times More Power Per Query Than Chatbots, KAIST Study Finds - finance.biggo.comnewsShow HN: Benchmark your eng team's AI agent maturity in 5 minutesnewsOpenAI is shutting down Atlas, but its AI browser ambitions are still growingnewsWeblica: Scalable and Reproducible Training Environments for Visual Web Agents - Apple Machine Learning ResearchnewsMeet WebBrain: An Open-Source, Local-First AI Browser Agent That Reads Pages and Automates Tasks in Chrome and Firefox - MarkTechPostnewsWhat AI benchmarks are not telling younewsFun Game - https://numdle-game.pages.dev/ [P]news10 Powerful Open-Source AI Tools for Local Hosting and Faster AI Queries - Geeky GadgetspaperEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentsnewsA randomized trial by METR found that experienced developers completed real coding tasks 19% slower when allowed to use AI tools — yet afterwards, they estimated on average that AI had made them 20% faster. - ScienceBlog.comnewsScarfBench: Benchmarking AI Agents for Enterprise Java Framework MigrationpaperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task SpecificationsnewsGoogle Cloud tests AI agents with ambiguity-based benchmarks - IT Brief AustralianewsAI Agents Consume 136 Times More Power Than Traditional Chatbots - Aju PressnewsCueBench for Developers is live: score how well you drive coding agentsnewsIntroducing GeneBench-PronewsDetecting silent agent failures with Amazon Bedrock AgentCore optimization

Topics