Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine narrowing but provides no primitive for fine-to-coarse backtracking. As a result, these agents typically converge prematurely and cannot recover from an early mistake. We propose VideoTreeSearch (VTS), a framework that casts grounded LVQA as iterative self-correcting search over an ad
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%We're building agents that can read millions of documents, but still forget a video they watched yesterday. →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “Searching Videos as Trees: Self-Correcting Agents for Ground” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “Searching Videos as Trees: Self-Correcting Agents for Ground” ≈ “Unity-Technologies/ml-agents””
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%Eigenwise/atomic-agents →
“Fuzzy title match (0.73): “Searching Videos as Trees: Self-Correcting Agents for Ground” ≈ “Eigenwise/atomic-agents””
- LinkedLinked via arxiv author · 85%Ce Zhang →
“Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA”
- LinkedLinked via arxiv author · 85%Ziyang Wang →
“Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA”
