COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COM
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%BerriAI/litellm →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: luo”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “COMET: Contrastive Motion-Enhanced Temporal Reasoning for Vi” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Chenghua Zhu →
“COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models”
- LinkedLinked via arxiv author · 85%Zhaolu Kang →
“COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models”
- LinkedLinked via arxiv author · 85%Qifan Shi →
“COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models”
- LinkedLinked via arxiv author · 85%Siyan Wu →
“COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models”
