CromSS: Cross-modal pre-training with noisy labels for remote sensing image segmentation
We study the potential of noisy labels y to pretrain semantic segmentation models in a multi-modal learning framework for geospatial applications. Specifically, we propose a novel Cross-modal Sample Selection method (CromSS) that utilizes the class distributions P^{(d)}(x,c) over pixels x and classes c modelled by multiple sensors/modalities d of a given geospatial scene. Consistency of predictions across sensors $d$ is jointly informed by the entropy of P^{(d)}(x,c). Noisy label sampling we determine by the confidence of each sensor d in the noisy class label, P^{(d)}(x,c=y(x)). To verify the performance of our approach, we conduct experiments with Sentinel-1 (radar) and Sentinel-2 (optical) satellite imagery from the globally-sampled SSL4EO-S12 dataset. We pair those scenes with 9-class noisy labels sourced from the Google Dynamic World project for pretraining. Transfer learning evaluations (downstream task) on the DFC2020 dataset confirm the effectiveness of the proposed method for remote sensing image segmentation.
Code (0)
등록된 구현이 없습니다.
Tasks
Image SegmentationSemantic SegmentationTransfer LearningSimilar Papers 제목 키워드 기반
Learning with Noisy Correspondence for Cross-modal Matching
Cross-modal matching, which aims to establish the correspondence between two different modalities, is fundamental to a variety of tasks such as cross-modal retrieval and vision-and-language understanding. Although a huge…
Cross-Modal RetrievalCross-modal retrieval with noisy correspondenceImage-text matchingMemorization+2A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
Cross-modal retrieval (CMR) aims to establish interaction between different modalities, among which supervised CMR is emerging due to its flexibility in learning semantic category discrimination. Despite the remarkable p…
Cross-Modal RetrievalRetrievalLearning Cross-Modal Retrieval With Noisy Labels
Recently, cross-modal retrieval is emerging with the help of deep multimodal learning. However, even for unimodal data, collecting large-scale well-annotated data is expensive and time-consuming, and not to mention t…
Cross-Modal RetrievalRetrievalMOLAR: Learning Multimodal Molecular Representations from Noisy Labels
Motivation: Noisy labels are a common challenge in molecular property prediction because molecular annotations are often obtained from assays, curated databases, or weak annotation pipelines rather than directly observed…
Molecular Property PredictionRepresentation LearningNoise-Tolerant Learning for Audio-Visual Action Recognition
Recently, video recognition is emerging with the help of multi-modal learning, which focuses on integrating distinct modalities to improve the performance or robustness of the model. Although various multi-modal learning…
Action RecognitionNoise EstimationVideo Recognition