UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures. To address these limitations, we introduce UniC
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 66%xlang-ai/OSWorld →
- PossiblePossibly related (embedding) · 61%TIGER-AI-Lab/ClawBench →
- PossiblePossibly related (embedding) · 60%zapier/AutomationBench →
- PossiblePossibly related (embedding) · 59%opencmit/alphora →
- PossiblePossibly related (embedding) · 58%AgentCore-8B →
- PossiblePossibly related (embedding) · 26%2FastLabs/agent-squad →
“Possibly related via embedding similarity 0.57 (not asserted). Timestamp check: artifact slightly before paper (-7d).”
- LinkedLinked via arxiv author · 85%Zhekai Chen →
“UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks”
- LinkedLinked via arxiv author · 85%Chengqi Duan →
“UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks”
