Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Human-JEPA: A Human-Centric Vision Model that Perceives and ” ≈ “VioletVision-3B””
- PossiblePossibly related (embedding) · 53%Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning →
- PossiblePossibly related (embedding) · 50%Why first person video may matter for robot learning[D] →
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “Human-JEPA: A Human-Centric Vision Model that Perceives and ” ≈ “pytorch/vision””
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Zhihui Wei →
“Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates”
- LinkedLinked via arxiv author · 85%Licai Sun →
“Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates”
- LinkedLinked via arxiv author · 85%Guoying Zhao →
“Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates”
