MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces diverse, visually grounded scenarios through information-flow-guided planning and multi-stage quali
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%sierra-research/tau2-bench →
- PossiblePossibly related (embedding) · 59%AgentCore-8B →
- PossiblePossibly related (embedding) · 59%xlang-ai/OSWorld →
- PossiblePossibly related (embedding) · 55%Pluggable by design: An agent mesh for software modernization that adopts the next model release →
- PossiblePossibly related (embedding) · 54%dustland/agentok →
- LinkedLinked via arxiv author · 85%Kaixin Ma →
“MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents”
- LinkedLinked via arxiv author · 85%Di Feng →
“MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents”
- LinkedLinked via arxiv author · 85%Alexander Metz →
“MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents”
