More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at su
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Stefan Maria Ailuro →
“More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe”
- LinkedLinked via arxiv author · 85%Mario Markov →
“More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe”
- LinkedLinked via arxiv author · 85%Mohammad Mahdi →
“More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe”
- LinkedLinked via arxiv author · 85%Luc Van Gool →
“More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe”
- LinkedLinked via arxiv author · 85%Danda Pani Paudel →
“More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe”
