A Vision Centric Remote Sensing Benchmark
Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents unique challenges that current MLLMs struggle to handle, particularly in visual grounding and spatial reasoning. This study investigates the limitations of CLIP-based MLLMs in RS, highlighting their failure to differentiate visually distinct yet semantically similar RS images. To address this, we introduce a remote sensing multimodal visual patterns (RSMMVP) benchmark. It is designed to evaluate MLLMs in RS tasks by identifying the CLIP-blind pairs, where CLIP-based models incorrectly assign high similarity scores to visually distinct RS images. Through a visual question answering (VQA) evaluation, we analyze the performance of state-of-the-art MLLMs, revealing significant limitations in RS specific representation learning. The results provide valuable insights into the weaknesses of CLIP-based visual encoding and offer a foundation for future research to develop more effective MLLMs tailored for remote sensing applications.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringRepresentation LearningSpatial ReasoningVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
An assessment of data-centric methods for label noise identification in remote sensing data sets
Label noise in the sense of incorrect labels is present in many real-world data sets and is known to severely limit the generalizability of deep learning models. In the field of remote sensing, however, automated treatme…
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been pro…
Image CaptioningQuestion AnsweringVisual GroundingVisual Question AnsweringEagleVision: Object-level Attribute Multimodal LLM for Remote Sensing
Recent advances in multimodal large language models (MLLMs) have demonstrated impressive results in various visual tasks. However, in remote sensing (RS), high resolution and small proportion of objects pose challenges t…
AttributeDisentanglementObjectobject-detection+1Self-supervised Learning in Remote Sensing: A Review
In deep learning research, self-supervised learning (SSL) has received great attention triggering interest within both the computer vision and remote sensing communities. While there has been a big success in computer vi…
Earth ObservationImage ClassificationMulti-Label Image ClassificationSelf-Supervised LearningLeveraging feature communication in federated learning for remote sensing image classification
In the realm of Federated Learning (FL) applied to remote sensing image classification, this study introduces and assesses several innovative communication strategies. Our exploration includes feature-centric communicati…
ClassificationFederated Learningimage-classificationImage Classification+2