VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation
We propose VT-MUSE, a Multimodal Unified SEquential representation learning framework for visuotactilemanipulation. Existing approaches often encode visual and tactile observations independently before fusion, limiting their ability to capture fine-grained cross-modal dependencies. Moreover, most methods focus on observations at the current time step and overlook the temporal evolution of contact. VT-MUSE addresses both limitations through a two-stage representation learning framework. In Stage I, modality specific encoders are jointly adapted via cross-modal temporal alignment and masked-view
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Introducing Gemma 4 12B: a unified, encoder-free multimodal model →
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%janhq/jan →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “VT-MUSE: Multimodal Unified Sequential Visuotactile Represen” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Congsheng Xu →
“VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation”
- LinkedLinked via arxiv author · 85%Qiaochu Yang →
“VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation”
