Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo →
“Fuzzy title match (0.73): “Multimodal Large Language Models for Remote Sensing Image Un” ≈ “Tongyi-MAI/Z-Image-Turbo””
- PossiblePossibly related (embedding) · 52%Embed the world: Multimodal AI for searchable aerial imagery at scale →
- FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail →
“Shared author/contributor keys: cheng”
- LinkedLinked via arxiv author · 85%Qiwei Ma →
“Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?”
- LinkedLinked via arxiv author · 85%Chunping Qiu →
“Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?”
- LinkedLinked via arxiv author · 85%Xinjun Cheng →
“Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?”
- LinkedLinked via arxiv author · 85%Xiaoyu Zhang →
“Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?”
- LinkedLinked via arxiv author · 85%Puhong Duan →
“Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?”
