VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-drive
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization - Apple Machine Learning Research →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “VideoChat3: Fully Open Video MLLM for Efficient and Generali” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Xinhao Li →
“VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding”
- LinkedLinked via arxiv author · 85%Yuhan Zhu →
“VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding”
- LinkedLinked via arxiv author · 85%Xiangyu Zeng →
“VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding”
