Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal reasoning may therefore introduce substantial redundancy and obscure decisive evidence. This motivates decoupling evidence acquisition from verification:
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%We're building agents that can read millions of documents, but still forget a video they watched yesterday. →
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%Fosowl/agenticSeek →
“Fuzzy title match (0.73): “Sparse Evidence Can Suffice: Agentic Evidence Seeking for Mu” ≈ “Fosowl/agenticSeek””
- LinkedLinked via arxiv author · 85%Haochen Zhao →
“Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection”
- LinkedLinked via arxiv author · 85%Yongxiu Xu →
“Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection”
