Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover uses each tool result to select the next action, so lo
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 55%We're building agents that can read millions of documents, but still forget a video they watched yesterday. →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “Thinking Beyond Videos: Unifying Video Reasoning and Deep Re” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “Thinking Beyond Videos: Unifying Video Reasoning and Deep Re” ≈ “Unity-Technologies/ml-agents””
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%rasbt/LLMs-from-scratch →
“Shared author/contributor keys: yin”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Wenqi Liu →
“Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents”
