paper-with-me

홈 › Papers

Context Patch Fusion With Class Token Enhancement for Weakly Supervised Semantic Segmentation

2026-01-21 · Yiyang Fu, Hui Li, Wangyu Wu arxiv

Weakly Supervised Semantic Segmentation (WSSS), which relies only on image-level labels, has attracted significant attention for its cost-effectiveness and scalability. Existing methods mainly enhance inter-class distinctions and employ data augmentation to mitigate semantic ambiguity and reduce spurious activations. However, they often neglect the complex contextual dependencies among image patches, resulting in incomplete local representations and limited segmentation accuracy. To address these issues, we propose the Context Patch Fusion with Class Token Enhancement (CPF-CTE) framework, which exploits contextual relations among patches to enrich feature representations and improve segmentation. At its core, the Contextual-Fusion Bidirectional Long Short-Term Memory (CF-BiLSTM) module captures spatial dependencies between patches and enables bidirectional information flow, yielding a more comprehensive understanding of spatial correlations. This strengthens feature learning and segmentation robustness. Moreover, we introduce learnable class tokens that dynamically encode and refine class-specific semantics, enhancing discriminative capability. By effectively integrating spatial and semantic cues, CPF-CTE produces richer and more accurate representations of image content. Extensive experiments on PASCAL VOC 2012 and MS COCO 2014 validate that CPF-CTE consistently surpasses prior WSSS methods.

📄 PDF Abstract BibTeX arXiv:2601.14718

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationData Augmentation

Similar Papers 제목 키워드 기반

Ultra-high Resolution Image Segmentation via Locality-aware Context Fusion and Alternating Local Enhancement

2021-09-06 · ICCV 2021 10 · Wenxi Liu, Qi Li, Xindai Lin, Weixiang Yang 외

Ultra-high resolution image segmentation has raised increasing interests in recent years due to its realistic applications. In this paper, we innovate the widely used high-resolution image segmentation pipeline, in which…

Image SegmentationLand Cover ClassificationSegmentationSemantic Segmentation

CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization

2026-03-08 · Anh-Duy Le, Van-Linh Pham, Thanh-Nam Vo, Xuan Toan Mai 외 arxiv

One-shot styled handwriting image generation, despite achieving impressive results in recent years, remains challenging due to the difficulty in capturing the intricate and diverse characteristics of human handwriting by…

Image Generation

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

2026-04-16 · Mingqian Ji, Shanshan Zhang, Jian Yang arxiv

Vision Transformer (ViT)-based sparse multi-view 3D object detectors have achieved remarkable accuracy but still suffer from high inference latency due to heavy token processing. To accelerate these models, token compres…

MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model

2026-03-27 · Quan Dao, Dimitris Metaxas arxiv

Transformer architectures, particularly Diffusion Transformers (DiTs), have become widely used in diffusion and flow-matching models due to their strong performance compared to convolutional UNets. However, the isotropic…

Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization

2024-06-13 · Qihao Liu, Zhanpeng Zeng, Ju He, Qihang Yu 외

This paper presents innovative enhancements to diffusion models by integrating a novel multi-resolution network and time-dependent layer normalization. Diffusion models have gained prominence for their effectiveness in h…

Image Generation