Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes. We first ask a diagnostic question: if the policy is given an expert-trajectory future image as privileged input at training and testing time, is that additional visual evidence useful for choosing the current action? (These expert-trajectory future images are unavailable at test time in real depl
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Anticipate Before Acting: Future-State-Conditioned Vision-La” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “Anticipate Before Acting: Future-State-Conditioned Vision-La” ≈ “pytorch/vision””
- LinkedLinked via arxiv author · 85%Lingfeng Zhang →
“Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation”
- LinkedLinked via arxiv author · 85%Zhanguang Zhang →
“Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation”
- LinkedLinked via arxiv author · 85%Liheng Ma →
“Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation”
- LinkedLinked via arxiv author · 85%Tongtong Cao →
“Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation”
- LinkedLinked via arxiv author · 85%Yingxue Zhang →
“Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation”
