MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation
Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored. We introduce MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, MonoIR-RS retains the infra
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%vlm-starter →
- PossiblePossibly related (embedding) · 45%Climate-Vision/ClimateVision →
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “MonoIR-RS: Infrared Remote Sensing Vision-Language Learning ” ≈ “VioletVision-3B””
- LinkedLinked via arxiv author · 85%Jiaju Han →
“MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation”
- LinkedLinked via arxiv author · 85%Ma Yaqi →
“MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation”
- LinkedLinked via arxiv author · 85%Yahui Chai →
“MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation”
- LinkedLinked via arxiv author · 85%Xuemeng Sun →
“MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation”
- LinkedLinked via arxiv author · 85%Xin Lin →
“MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation”
