VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-drive
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 49%Laion Big Video Dataset →
- PossiblePossibly related (embedding) · 52%Getting video models to learn better, faster →
- PossiblePossibly related (embedding) · 51%Introducing agentic video understanding with Gemini →
- PossiblePossibly related (embedding) · 51%Machine learning for 360° video streaming over 5G and B5G networks: emerging applications and challenges - Springer Nature Link →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
