Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, thes
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Lars Benedikt Kaesberg →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
- LinkedLinked via arxiv author · 85%Tianyu Yang →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
- LinkedLinked via arxiv author · 85%Florian Valentin Wunderlich →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
- LinkedLinked via arxiv author · 85%Terry Ruas →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
- LinkedLinked via arxiv author · 85%Jan Philip Wahle →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
- LinkedLinked via arxiv author · 85%Daniel Kurzawe →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
- LinkedLinked via arxiv author · 85%Bela Gipp →
“Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds”
