IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generat
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%janhq/jan →
“Shared author/contributor keys: han”
- PossiblePossibly related (embedding) · 53%Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning - Apple Machine Learning Research →
- LinkedLinked via arxiv author · 85%Jiapeng Li →
“IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning”
- LinkedLinked via arxiv author · 85%Ping Wei →
“IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning”
- LinkedLinked via arxiv author · 85%Wenjuan Han →
“IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning”
- LinkedLinked via arxiv author · 85%Song-chun Zhu →
“IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning”
- LinkedLinked via arxiv author · 85%Lifeng Fan →
“IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning”
