paper-with-me

홈 › Papers

SeMask: Semantically Masked Transformers for Semantic Segmentation

2021-12-23 · arXiv 2021 12 · Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, Humphrey Shi

Finetuning a pretrained backbone in the encoder part of an image transformer network has been the traditional approach for the semantic segmentation task. However, such an approach leaves out the semantic context that an image provides during the encoding stage. This paper argues that incorporating semantic information of the image into pretrained hierarchical transformer-based backbones while finetuning improves the performance considerably. To achieve this, we propose SeMask, a simple and effective framework that incorporates semantic information into the encoder with the help of a semantic attention operation. In addition, we use a lightweight semantic decoder during training to provide supervision to the intermediate semantic prior maps at every stage. Our experiments demonstrate that incorporating semantic priors enhances the performance of the established hierarchical encoders with a slight increase in the number of FLOPs. We provide empirical proof by integrating SeMask into Swin Transformer and Mix Transformer backbones as our encoder paired with different decoders. Our framework achieves a new state-of-the-art of 58.25% mIoU on the ADE20K dataset and improvements of over 3% in the mIoU metric on the Cityscapes dataset. The code and checkpoints are publicly available at https://github.com/Picsart-AI-Research/SeMask-Segmentation .

📄 PDF Abstract BibTeX arXiv:2112.12782

Code (1)

Picsart-AI-Research/SeMask-Segmentation 공식 구현 pytorch

Tasks

DecoderSemantic Segmentation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Image BERT Pre-training with Online Tokenizer

2021-09-29 · ICLR 2022 4 · Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen 외

The success of language Transformers is primarily attributed to the pretext task of masked language modeling (MLM), where texts are first tokenized into semantically meaningful pieces. In this work, we study masked image…

image-classificationImage ClassificationInstance SegmentationLanguage Modeling+5

iBOT: Image BERT Pre-Training with Online Tokenizer

2021-11-15 · Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen 외

The success of language Transformers is primarily attributed to the pretext task of masked language modeling (MLM), where texts are first tokenized into semantically meaningful pieces. In this work, we study masked image…

image-classificationImage ClassificationInstance SegmentationLanguage Modeling+7

Masked Clustering Prediction for Unsupervised Point Cloud Pre-training

2025-08-12 · Bin Ren, Xiaoshui Huang, Mengyuan Liu, Hong Liu 외 arxiv

Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative se…

Unsupervised Pre-trainingSemantic SegmentationContrastive LearningObject Detection

SQ-GAN: Semantic Image Communications Using Masked Vector Quantization

2025-02-13 · Francesco Pezone, Sergio Barbarossa, Giuseppe Caire

This work introduces Semantically Masked VQ-GAN (SQ-GAN), a novel approach integrating generative models to optimize image compression for semantic/task-oriented communications. SQ-GAN employs off-the-shelf semantic sema…

Image CompressionQuantizationSegmentationSemantic Segmentation

DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture

2024-05-28 · Shentong Mo, Sukmin Yun

The joint-embedding predictive architecture (JEPA) recently has shown impressive results in extracting visual representations from unlabeled imagery under a masking strategy. However, we reveal its disadvantages, notably…

image-classificationImage Classificationobject-detectionObject Detection+1