paper-with-me

홈 › Papers

SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting

2024-12-11 · Pallavi Jain, Dino Ienco, Roberto Interdonato, Tristan Berchoux, Diego Marcos

Pre-trained vision-language models (VLMs), such as CLIP, demonstrate impressive zero-shot classification capabilities with free-form prompts and even show some generalization in specialized domains. However, their performance on satellite imagery is limited due to the underrepresentation of such data in their training sets, which predominantly consist of ground-level images. Existing prompting techniques for satellite imagery are often restricted to generic phrases like a satellite image of ..., limiting their effectiveness for zero-shot land-use and land-cover (LULC) mapping. To address these challenges, we introduce SenCLIP, which transfers CLIPs representation to Sentinel-2 imagery by leveraging a large dataset of Sentinel-2 images paired with geotagged ground-level photos from across Europe. We evaluate SenCLIP alongside other SOTA remote sensing VLMs on zero-shot LULC mapping tasks using the EuroSAT and BigEarthNet datasets with both aerial and ground-level prompting styles. Our approach, which aligns ground-level representations with satellite imagery, demonstrates significant improvements in classification accuracy across both prompt styles, opening new possibilities for applying free-form textual descriptions in zero-shot LULC mapping.

📄 PDF Abstract BibTeX arXiv:2412.08536

Code (1)

pallavijain-pj/SenCLIP 공식 구현 pytorch

Tasks

zero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing

2025-08-16 · Pallavi Jain, Diego Marcos, Dino Ienco, Roberto Interdonato 외 arxiv

Vision-language models (VLMs) have shown significant promise in remote sensing applications, particularly for land-use and land-cover (LULC) mapping via zero-shot classification and retrieval. However, current approaches…

LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping

2025-11-11 · Chenying Liu, Wei Huang, Xiao Xiang Zhu arxiv

Land Use and Land Cover (LULC) mapping is a fundamental task in Earth Observation (EO). However, current LULC models are typically developed for a specific modality and a fixed class taxonomy, limiting their generability…

Transfer Learning

Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment

2025-08-27 · Namu Kim, Wonbin Kweon, Minsoo Kim, Hwanjo Yu arxiv

We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by t…

Image Generation

DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications

2025-02-24 · Ibrahim Fayad, Max Zimmer, Martin Schwartz, Philippe Ciais 외

Significant efforts have been directed towards adapting self-supervised multimodal learning for Earth observation applications. However, existing methods produce coarse patch-sized embeddings, limiting their effectivenes…

cross-modal alignmentEarth Observation

Generalized Few-Shot Meets Remote Sensing: Discovering Novel Classes in Land Cover Mapping via Hybrid Semantic Segmentation Framework

2024-04-19 · Zhuohong Li, Fangxiao Lu, Jiaqi Zou, Lei Hu 외

Land-cover mapping is one of the vital applications in Earth observation, aiming at classifying each pixel's land-cover type of remote-sensing images. As natural and human activities change the landscape, the land-cover …

Earth ObservationSegmentationSemantic Segmentation