AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world applications in specialized fields, where models inevitably encounter rare visual concepts and complex spatio-temporal dynamics. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adap
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%vlm-starter →
- PossiblePossibly related (embedding) · 53%VioletVision-3B →
- PossiblePossibly related (embedding) · 50%open-edge-platform/geti →
- LinkedLinked via arxiv author · 85%Rintaro Otsubo →
“AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models”
- LinkedLinked via arxiv author · 85%Ryo Fujii →
“AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models”
- LinkedLinked via arxiv author · 85%Reina Ishikawa →
“AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models”
- LinkedLinked via arxiv author · 85%Taiki Kanaya →
“AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models”
- LinkedLinked via arxiv author · 85%Kanta Sawafuji →
“AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models”
