GeoWAM: Visual Geometry World Action Models for Autonomous Driving
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action →
“Fuzzy title match (0.92): “GeoWAM: Visual Geometry World Action Models for Autonomous D” ≈ “liguodongiot/llm-action””
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%mudler/LocalAI →
“Shared author/contributor keys: guo”
- FuzzyOverlapping authors or contributors · 62%rasbt/LLMs-from-scratch →
“Shared author/contributor keys: yin”
- LinkedLinked via arxiv author · 85%Yiren Lu →
“GeoWAM: Visual Geometry World Action Models for Autonomous Driving”
- LinkedLinked via arxiv author · 85%Xin Ye →
“GeoWAM: Visual Geometry World Action Models for Autonomous Driving”
- LinkedLinked via arxiv author · 85%Jiaming Liu →
“GeoWAM: Visual Geometry World Action Models for Autonomous Driving”
- LinkedLinked via arxiv author · 85%Jin Yao →
“GeoWAM: Visual Geometry World Action Models for Autonomous Driving”
