Learning Content-enhanced Mask Transformer for Domain Generalized Urban-Scene Segmentation
Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike domain gap challenges, USSS is unique in that the semantic categories are often similar in different urban scenes, while the styles can vary significantly due to changes in urban landscapes, weather conditions, lighting, and other factors. Existing approaches typically rely on convolutional neural networks (CNNs) to learn the content of urban scenes. In this paper, we propose a Content-enhanced Mask TransFormer (CMFormer) for domain-generalized USSS. The main idea is to enhance the focus of the fundamental component, the mask attention mechanism, in Transformer segmentation models on content information. To achieve this, we introduce a novel content-enhanced mask attention mechanism. It learns mask queries from both the image feature and its down-sampled counterpart, as lower-resolution image features usually contain more robust content information and are less sensitive to style variations. These features are fused into a Transformer decoder and integrated into a multi-resolution content-enhanced mask attention learning scheme. Extensive experiments conducted on various domain-generalized urban-scene segmentation datasets demonstrate that the proposed CMFormer significantly outperforms existing CNN-based methods for domain-generalized semantic segmentation, achieving improvements of up to 14.00\% in terms of mIoU (mean intersection over union). The source code is publicly available at \url{https://github.com/BiQiWHU/CMFormer}.
Code (1)
Tasks
DecoderDomain AdaptationDomain GeneralizationScene SegmentationSegmentationSemantic SegmentationSource-Free Domain AdaptationSynthetic-to-Real TranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Generalized Segmentation for Foggy-scenes by Bi-directional Wavelet Guidance
Learning scene semantics that can be well generalized to foggy conditions is important for safety-crucial applications such as autonomous driving. Existing methods need both annotated clear images and foggy images to tr…
Autonomous DrivingDomain AdaptationDomain GeneralizationFoggy Scene Segmentation+3TSTNN: Two-stage Transformer based Neural Network for Speech Enhancement in the Time Domain
In this paper, we propose a transformer-based architecture, called two-stage transformer neural network (TSTNN) for end-to-end speech denoising in the time domain. The proposed model is composed of an encoder, a two-stag…
DecoderDenoisingSpeech DenoisingSpeech EnhancementTextual Query-Driven Mask Transformer for Domain Generalized Segmentation
In this paper, we introduce a method to tackle Domain Generalized Semantic Segmentation (DGSS) by utilizing domain-invariant semantic knowledge from text embeddings of vision-language models. We employ the text embedding…
Domain GeneralizationObjectSemantic SegmentationHGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation
Current semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain tha…
Domain GeneralizationSegmentationSemantic SegmentationAttention Head Masking for Inference Time Content Selection in Abstractive Summarization
How can we effectively inform content selection in Transformer-based abstractive summarization models? In this work, we present a simple-yet-effective attention head masking technique, which is applied on encoder-decoder…
Abstractive Text SummarizationDecoderDocument Summarization