paper-with-me

Papers

Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis

2026-04-01 · Xingxing Weng, Ruifeng Ni, Chao Pang, XiangYu Hao, Yishan Wang, Xiaokang Zhang, Wei Xu, Gui-Song Xia arxiv

Current remote sensing vision-language models (RS VLMs) demonstrate impressive performance in image interpretation but rely on static training data, limiting their ability to accommodate continuously emerging sensing modalities and downstream tasks. This exposes a fundamental challenge: enabling RS VLMs to continually adapt without catastrophic forgetting. Despite its practical importance, the continual learning capability of RS VLMs remains underexplored, and no dedicated benchmark currently exists. In this work, we present CLeaRS, a comprehensive benchmark for continual vision-language learning in remote sensing. CLeaRS comprises 10 curated subsets with over 207k image-text pairs, spanning diverse interpretation tasks, sensing modalities, and application scenarios. We further define three evaluation protocols: long-horizon, modality-incremental, and task-incremental settings, to systematically assess continual adaptation. Extensive benchmarking of diverse vision-language models reveals catastrophic forgetting across all settings. Moreover, representative continual learning methods, when adapted to RS VLMs, exhibit limited effectiveness in handling task, instruction, and modality transitions. Our findings underscore the need for developing continual learning methods tailored to RS VLMs.

📄 PDF Abstract BibTeX arXiv:2604.00820

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing

2023-12-20 · Zhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu 외

Remote sensing imagery, despite its broad applications in helping achieve Sustainable Development Goals and tackle climate change, has not yet benefited from the recent advancements of versatile, task-agnostic vision lan…

AttributeCross-Modal RetrievalImage GenerationRetrieval+1

RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models

2024-08-27 · Junyao Ge, Xu Zhang, Yang Zheng, Kaitai Guo 외

Abundant, well-annotated multimodal data in remote sensing are pivotal for aligning complex visual remote sensing (RS) scenes with human language, enabling the development of specialized vision language models across div…

DescriptiveLanguage ModelingLanguage ModellingScene Understanding

CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models

2024-11-27 · Xiao An, Jiaxing Sun, Zihan Gui, wei he

The rapid advancement of Large Vision-Language Models (VLMs), both general-domain models and those specifically tailored for remote sensing, has demonstrated exceptional perception and reasoning capabilities in Earth obs…

BenchmarkingEarth ObservationMultiple-choice

Weakly-supervised continual learning for class-incremental segmentation

2022-01-04 · Gaston Lenczner, Adrien Chan-Hon-Tong, Nicola Luminari, Bertrand Le Saux

Transfer learning is a powerful way to adapt existing deep learning models to new emerging use-cases in remote sensing. Starting from a neural network already trained for semantic segmentation, we propose to modify its l…

Continual LearningPseudo LabelSemantic SegmentationTransfer Learning

Remote Sensing Image Classification with the SEN12MS Dataset

2021-04-01 · Michael Schmitt, Yu-Lun Wu

Image classification is one of the main drivers of the rapid developments in deep learning with convolutional neural networks for computer vision. So is the analogous task of scene classification in remote sensing. Howev…

BenchmarkingClassificationGeneral Classificationimage-classification+3