paper-with-me

Papers

Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP

2024-12-27 · Zhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su, Zhiyi Zhao, Ge Zhang, Jionglong Su, ZongYuan Ge

The application of Contrastive Language-Image Pre-training (CLIP) in Weakly Supervised Semantic Segmentation (WSSS) research powerful cross-modal semantic understanding capabilities. Existing methods attempt to optimize input text prompts for improved alignment of images and text, by finely adjusting text prototypes to facilitate semantic matching. Nevertheless, given the modality gap between text and vision spaces, the text prototypes employed by these methods have not effectively established a close correspondence with pixel-level vision features. In this work, our theoretical analysis indicates that the inherent modality gap results in misalignment of text and region features, and that this gap cannot be sufficiently reduced by minimizing contrast loss in CLIP. To mitigate the impact of the modality gap, we propose a Vision Prototype Learning (VPL) framework, by introducing more representative vision prototypes. The core of this framework is to learn class-specific vision prototypes in vision space with the help of text prototypes, for capturing high-quality localization maps. Moreover, we propose a regional semantic contrast module that contrasts regions embedding with corresponding prototypes, leading to more comprehensive and robust feature learning. Experimental results show that our proposed framework achieves state-of-the-art performance on two benchmark datasets.

📄 PDF Abstract BibTeX arXiv:2412.19650

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

DualProtoSeg: Simple and Efficient Design with Text- and Image-Guided Prototype Learning for Weakly Supervised Histopathology Image Segmentation

2025-12-11 · Anh M. Vu, Khang P. Le, Trang T. K. Vo, Ha Thach 외 arxiv

Weakly supervised semantic segmentation (WSSS) in histopathology seeks to reduce annotation cost by learning from image-level labels, yet it remains limited by inter-class homogeneity, intra-class heterogeneity, and the …

Semantic SegmentationImage Segmentation

Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation

2024-03-12 · CVPR 2024 1 · Feilong Tang, Zhongxing Xu, Zhaojun Qu, Wei Feng 외

Recent weakly supervised semantic segmentation (WSSS) methods strive to incorporate contextual knowledge to improve the completeness of class activation maps (CAM). In this work, we argue that the knowledge bias between …

Learning TheorySemantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Cross-Modal Prototype Allocation: Unsupervised Slide Representation Learning via Patch-Text Contrast in Computational Pathology

2025-03-26 · Yuxuan Chen, Jiawen Li, Jiali Hu, Xitong Ling 외

With the rapid advancement of pathology foundation models (FMs), the representation learning of whole slide images (WSIs) attracts increasing attention. Existing studies develop high-quality patch feature extractors and …

DescriptiveLarge Language ModelMultiple Instance LearningRepresentation Learning+1

ConStruct: Structural Distillation of Foundation Models for Prototype-Based Weakly Supervised Histopathology Segmentation

2025-12-11 · Khang Le, Ha Thach, Anh M. Vu, Trang T. K. Vo 외 arxiv

Weakly supervised semantic segmentation (WSSS) in histopathology relies heavily on classification backbones, yet these models often localize only the most discriminative regions and struggle to capture the full spatial e…

Semantic Segmentation

Weakly Supervised 3D Point Cloud Segmentation via Multi-Prototype Learning

2022-05-06 · Yongyi Su, Xun Xu, Kui Jia

Addressing the annotation challenge in 3D Point Cloud segmentation has inspired research into weakly supervised learning. Existing approaches mainly focus on exploiting manifold and pseudo-labeling to make use of large u…

Point Cloud SegmentationSegmentationWeakly Supervised 3D Point Cloud SegmentationWeakly-supervised Learning