WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a world-model pretraining objective. We present WALDO, a one-shot exemplar- and language-conditioned detection head with 3.4M trainable parameters that reads frozen V-JEPA 2.1 features to jointly predict object localization and target presence, with no gradient on the backbone. Be
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%microsoft/ML-For-Beginners →
“Shared author/contributor keys: gupta”
- LinkedLinked via arxiv author · 85%Kishor Datta Gupta →
“WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes”
- LinkedLinked via arxiv author · 85%Ahmed Rafi Hasan →
“WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes”
- LinkedLinked via arxiv author · 85%Md. Mahfuzur Rahman →
“WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes”
- LinkedLinked via arxiv author · 85%Md. Sadman Haque →
“WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes”
- LinkedLinked via arxiv author · 85%Mohd Ariful Haque →
“WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes”
