Read original ↗
paperarXivTrust 82 · PrimaryPublished 10d agoLive · 7d ago

ID-VTG: Image-Disambiguated Video Temporal Grounding

Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone. To address this, we introduce Image-Disambiguated Video Temporal Grounding (ID-VTG), a task that leverages multimodal queries combining a reference image and a text description to precisely localize segments where a specific instance performs a described action. To facilitate research, we construct two benchmarks

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo

    Fuzzy title match (0.73): “ID-VTG: Image-Disambiguated Video Temporal Grounding” ≈ “Tongyi-MAI/Z-Image-Turbo”

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses

    Fuzzy title match (0.73): “ID-VTG: Image-Disambiguated Video Temporal Grounding” ≈ “Developer-Y/cs-video-courses”

  • LinkedLinked via arxiv author · 85%Minghang Zheng

    ID-VTG: Image-Disambiguated Video Temporal Grounding

  • LinkedLinked via arxiv author · 85%Jingli Wei

    ID-VTG: Image-Disambiguated Video Temporal Grounding

  • LinkedLinked via arxiv author · 85%Hongyi Yang

    ID-VTG: Image-Disambiguated Video Temporal Grounding

  • LinkedLinked via arxiv author · 85%Dongyang Liu

    ID-VTG: Image-Disambiguated Video Temporal Grounding

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics