Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reasoning from the appearance and dynamics of hands and objects themselves. To address this limitation, we propose a new learning paradigm that combines (i) hand-object masked training, which enables robust reasoning from partial hand or object observations, and (ii) an HOI-dynamics-aware decoder that explicitly learns hand
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Masatoshi Tateno →
“Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?”
- LinkedLinked via arxiv author · 85%Alexandros Stergiou →
“Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?”
- LinkedLinked via arxiv author · 85%Risa Shinoda →
“Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?”
- LinkedLinked via arxiv author · 85%Yoichi Sato →
“Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?”
- LinkedLinked via arxiv author · 85%Dima Damen →
“Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “Do Egocentric Video-Language Models Capture Both Hand- and O” ≈ “Developer-Y/cs-video-courses””
- PossiblePossibly related (embedding) · 59%Why first person video may matter for robot learning[D] →
