paper-with-me

Papers

CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset

2024-11-18 · Zhiming Wang, Mingze Wang, Sheng Xu, Yanjing Li, Baochang Zhang

Remote Sensing Image Change Captioning (RSICC) aims to generate natural language descriptions of surface changes between multi-temporal remote sensing images, detailing the categories, locations, and dynamics of changed objects (e.g., additions or disappearances). Many current methods attempt to leverage the long-sequence understanding and reasoning capabilities of multimodal large language models (MLLMs) for this task. However, without comprehensive data support, these approaches often alter the essential feature transmission pathways of MLLMs, disrupting the intrinsic knowledge within the models and limiting their potential in RSICC. In this paper, we propose a novel model, CCExpert, based on a new, advanced multimodal large model framework. Firstly, we design a difference-aware integration module to capture multi-scale differences between bi-temporal images and incorporate them into the original image context, thereby enhancing the signal-to-noise ratio of differential features. Secondly, we constructed a high-quality, diversified dataset called CC-Foundation, containing 200,000 image pairs and 1.2 million captions, to provide substantial data support for continue pretraining in this domain. Lastly, we employed a three-stage progressive training process to ensure the deep integration of the difference-aware integration module with the pretrained MLLM. CCExpert achieved a notable performance of $S^*_m=81.80$ on the LEVIR-CC benchmark, significantly surpassing previous state-of-the-art methods. The code and part of the dataset will soon be open-sourced at https://github.com/Meize0729/CCExpert.

📄 PDF Abstract BibTeX arXiv:2411.11360

Code (1)

meize0729/ccexpert 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

2026-07-24 · Yuheng Zong, Minghua Wang, Xin Zhao, Zhi-Hui Zhan 외 arxiv

Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation applications require fine-grained scenario specialization, constrained by scarce hig…

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

2026-07-22 · Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang 외 arxiv

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery…

Visual Question AnsweringRelational ReasoningScene UnderstandingVisual Grounding

From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing

2024-11-05 · Xintian Sun, Benji Peng, Charles Zhang, Fei Jin 외

Remote sensing has evolved from simple image acquisition to complex systems capable of integrating and processing visual and textual data. This review examines the development and application of multi-modal language mode…

Change DetectionContrastive LearningDisaster ResponseDomain Adaptation+7

VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing

2026-02-04 · Zhiming Luo, Di Wang, Haonan Guo, Jing Zhang 외 arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition a…

Scene ClassificationMultimodal ReasoningObject Recognition

BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model

2025-09-07 · Yujie Li, Wenjia Xu, Yuanben Zhang, Zhiwei Wei 외 arxiv

Bi-temporal satellite imagery supports critical applications such as urbanization monitoring and disaster assessment. Although powerful multimodal large language models~(MLLMs) have been applied in bi-temporal change ana…

Visual Question Answering