Read original ↗
paperarXivTrust 82 · PrimaryPublished 24d agoLive · 23d ago

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo

    Fuzzy title match (0.73): “Multimodal Large Language Models for Remote Sensing Image Un” ≈ “Tongyi-MAI/Z-Image-Turbo”

  • FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail

    Shared author/contributor keys: cheng

  • LinkedLinked via arxiv author · 85%Qiwei Ma

    Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

  • LinkedLinked via arxiv author · 85%Chunping Qiu

    Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

  • LinkedLinked via arxiv author · 85%Xinjun Cheng

    Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

  • LinkedLinked via arxiv author · 85%Xiaoyu Zhang

    Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

  • LinkedLinked via arxiv author · 85%Puhong Duan

    Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

Has model

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics