Whareformer: Learning to Track What is Where in Long Egocentric Videos
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become heavily occluded. In this paper, we propose the first learning-based solution to the OSNOM task: Whareformer, a transformer-based model with two components: an updatable memory of established tracks and a track assignment module that associates observations with existing tracks in a feed-forward manner. Whareformer join
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning →
- PossiblePossibly related (embedding) · 50%roboflow/supervision →
- PossiblePossibly related (embedding) · 50%Somnusochi/VLM-AutoYOLO →
- PossiblePossibly related (embedding) · 46%pytorch/vision →
- LinkedLinked via arxiv author · 85%Jacob Chalk →
“Whareformer: Learning to Track What is Where in Long Egocentric Videos”
- LinkedLinked via arxiv author · 85%Saptarshi Sinha →
“Whareformer: Learning to Track What is Where in Long Egocentric Videos”
- LinkedLinked via arxiv author · 85%Dima Damen →
“Whareformer: Learning to Track What is Where in Long Egocentric Videos”
- LinkedLinked via arxiv author · 85%Yannis Kalantidis →
“Whareformer: Learning to Track What is Where in Long Egocentric Videos”
