paper-with-me

홈 › Papers

MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation

2025-10-12 · Md Abdullah Al Mazid, Liangdong Deng, Naphtali Rishe arxiv

Clouds remain a major obstacle in optical satellite imaging, limiting accurate environmental and climate analysis. To address the strong spectral variability and the large scale differences among cloud types, we propose MSCloudCAM, a novel multi-scale context adapter network with convolution based cross-attention tailored for multispectral and multi-sensor cloud segmentation. A key contribution of MSCloudCAM is the explicit modeling of multiple complementary multi-scale context extractors. And also, rather than simply stacking or concatenating their outputs, our formulation uses one extractor's fine-resolution features and the other extractor's global contextual representations enabling dynamic, scale-aware feature selection. Building on this idea, we design a new convolution-based cross attention adapter that effectively fuses localized, detailed information with broader multi-scale context. Integrated with a hierarchical vision backbone and refined through channel and spatial attention mechanisms, MSCloudCAM achieves strong spectral-spatial discrimination. Experiments on various multisensor datatsets e.g. CloudSEN12 (Sentinel-2) and L8Biome (Landsat-8), demonstrate that MSCloudCAM achieves superior overall segmentation performance and competitive class-wise accuracy compared to recent state-of-the-art models, while maintaining competitive model complexity, highlighting the novelty and effectiveness of the proposed design for large-scale Earth observation.

📄 PDF Abstract BibTeX arXiv:2510.10802

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Scale Context Aggregation by Dilated Convolutions

2015-11-23 · Fisher Yu, Vladlen Koltun

State-of-the-art models for semantic segmentation are based on adaptations of convolutional networks that had originally been designed for image classification. However, dense prediction and image classification are stru…

General Classificationimage-classificationPredictionReal-Time Semantic Segmentation+2

CaseNet: Content-Adaptive Scale Interaction Networks for Scene Parsing

2019-04-17 · Xin Jin, Cuiling Lan, Wen-Jun Zeng, Zhizheng Zhang 외

Objects at different spatial positions in an image exhibit different scales. Adaptive receptive fields are expected to capture suitable ranges of context for accurate pixel level semantic prediction. Recently, atrous con…

PositionScene Parsing

Convolutional Neural Operators for robust and accurate learning of PDEs

2023-02-02 · NeurIPS 2023 11 · Bogdan Raonić, Roberto Molinaro, Tim De Ryck, Tobias Rohner 외

Although very successfully used in conventional machine learning, convolution based neural network architectures -- believed to be inconsistent in function space -- have been largely ignored in the context of learning so…

Operator learningPDE Surrogate Modeling

Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation

2021-09-08 · CVPR 2022 1 · Nan Xue, Tianfu Wu, Gui-Song Xia, Liangpei Zhang

This paper studies the problem of multi-person pose estimation in a bottom-up fashion. With a new and strong observation that the localization issue of the center-offset formulation can be remedied in a local-window sear…

Multi-Person Pose EstimationPose Estimation

Clay-CNN Hybrids: Leveraging Geospatial Foundation Models as Auxiliary Context for Landslide Detection

2026-06-12 · Huong Binh Vu arxiv

Rapid post-event landslide mapping is essential for disaster response but remains difficult to automate due to extreme class imbalance. This study evaluates whether Clay v1.5, a Geospatial Foundation Model (GFM), can imp…