Masked Cross-image Encoding for Few-shot Segmentation
Few-shot segmentation (FSS) is a dense prediction task that aims to infer the pixel-wise labels of unseen classes using only a limited number of annotated images. The key challenge in FSS is to classify the labels of query pixels using class prototypes learned from the few labeled support exemplars. Prior approaches to FSS have typically focused on learning class-wise descriptors independently from support images, thereby ignoring the rich contextual information and mutual dependencies among support-query features. To address this limitation, we propose a joint learning method termed Masked Cross-Image Encoding (MCE), which is designed to capture common visual properties that describe object details and to learn bidirectional inter-image dependencies that enhance feature interaction. MCE is more than a visual representation enrichment module; it also considers cross-image mutual dependencies and implicit guidance. Experiments on FSS benchmarks PASCAL-$5^i$ and COCO-$20^i$ demonstrate the advanced meta-learning ability of the proposed method.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot Semantic SegmentationSimilar Papers 제목 키워드 기반
R-MAE: Regions Meet Masked Autoencoders
In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked reg…
Contrastive LearningInteractive Segmentationobject-detectionObject Detection+2Self-Guided and Cross-Guided Learning for Few-Shot Segmentation
Few-shot segmentation has been attracting a lot of attention due to its effectiveness to segment unseen object classes with a few annotated samples. Most existing approaches use masked Global Average Pooling (GAP) to enc…
Few-Shot Semantic SegmentationImage SegmentationSegmentationSemantic SegmentationGeometry Aware Field-to-field Transformations for 3D Semantic Segmentation
We present a novel approach to perform 3D semantic segmentation solely from 2D supervision by leveraging Neural Radiance Fields (NeRFs). By extracting features along a surface point cloud, we achieve a compact representa…
3D Semantic SegmentationNeRFSegmentationSemantic SegmentationFLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
Remote sensing imagery is dense with objects and contextual visual information. There is a recent trend to combine paired satellite images and text captions for pretraining performant encoders for downstream tasks. Howev…
ClassificationContrastive LearningSemantic Segmentationzero-shot-classification+1A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our p…
Contrastive LearningDenoisingSemantic Segmentation