PercepCap: Video Captioner with Structured Spatio-Temporal Perception
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without exposing the perceptual evidence behind descriptions. As a result, mistakes in spatiotemporal perception are only observed in the final caption, making it difficult to identify the underlying perceptual errors directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evide
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: wan”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “PercepCap: Video Captioner with Structured Spatio-Temporal P” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Yifan Xu →
“PercepCap: Video Captioner with Structured Spatio-Temporal Perception”
- LinkedLinked via arxiv author · 85%Zihao Wang →
“PercepCap: Video Captioner with Structured Spatio-Temporal Perception”
- LinkedLinked via arxiv author · 85%Zhixiao Wang →
“PercepCap: Video Captioner with Structured Spatio-Temporal Perception”
- LinkedLinked via arxiv author · 85%Jiaming Zhang →
“PercepCap: Video Captioner with Structured Spatio-Temporal Perception”
