What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models
Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge across transformer layers, and how they are geometrically organized. In this work, we tackle these three questions through a systematic layer-wise analysis of V-JEPA 2 and VideoMAE-v2. We leverage lightweight probes trained to discover three temporally grounded properties: (i) camera motion understanding, (ii) intuitive physics, and (iii) anomaly detection. Both models encode camera motion, with best results ($>90$ ROC AU
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning →
- LinkedLinked via arxiv author · 85%Sharon S. Musa →
“What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models”
- LinkedLinked via arxiv author · 85%Fereshteh Forghani →
“What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models”
- LinkedLinked via arxiv author · 85%Harrish Thasarathan →
“What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models”
- LinkedLinked via arxiv author · 85%Sonia Joseph →
“What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models”
- LinkedLinked via arxiv author · 85%Matthew Kowal →
“What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models”
- LinkedLinked via arxiv author · 85%Konstantinos G. Derpanis →
“What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “What, Where, and How: Probing Spatiotemporal Representations” ≈ “Developer-Y/cs-video-courses””
