Read original ↗
paperarXivTrust 82 · PrimaryPublished 10d agoLive · 6d ago

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COM

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: luo

  • FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses

    Fuzzy title match (0.73): “COMET: Contrastive Motion-Enhanced Temporal Reasoning for Vi” ≈ “Developer-Y/cs-video-courses”

  • LinkedLinked via arxiv author · 85%Chenghua Zhu

    COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Zhaolu Kang

    COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Qifan Shi

    COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Siyan Wu

    COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

Implements (incoming)

authored (incoming)

Related across the graph

Topics