paper-with-me

홈 › Papers

DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation

2025-05-16 · CVPR 2025 1 · Ziyu Zhao, Xiaoguang Li, Linjia Shi, Nasrin Imanpour, Song Wang

Open-vocabulary semantic segmentation aims to segment images into distinct semantic regions for both seen and unseen categories at the pixel level. Current methods utilize text embeddings from pre-trained vision-language models like CLIP but struggle with the inherent domain gap between image and text embeddings, even after extensive alignment during training. Additionally, relying solely on deep text-aligned features limits shallow-level feature guidance, which is crucial for detecting small objects and fine details, ultimately reducing segmentation accuracy. To address these limitations, we propose a dual prompting framework, DPSeg, for this task. Our approach combines dual-prompt cost volume generation, a cost volume-guided decoder, and a semantic-guided prompt refinement strategy that leverages our dual prompting scheme to mitigate alignment issues in visual prompt generation. By incorporating visual embeddings from a visual prompt encoder, our approach reduces the domain gap between text and image embeddings while providing multi-level guidance through shallow features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches on multiple public datasets.

📄 PDF Abstract BibTeX arXiv:2505.11676

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

MEDPSeg: Hierarchical polymorphic multitask learning for the segmentation of ground-glass opacities, consolidation, and pulmonary structures on computed tomography

2023-12-04 · Diedre S. Carmo, Jean A. Ribeiro, Alejandro P. Comellas, Joseph M. Reinhardt 외

The COVID-19 pandemic response highlighted the potential of deep learning methods in facilitating the diagnosis, prognosis and understanding of lung diseases through automated segmentation of pulmonary structures and les…

AnatomyComputed Tomography (CT)Decision MakingPrognosis+1

Eff-3DPSeg: 3D organ-level plant shoot segmentation using annotation-efficient point clouds

2022-12-20 · Liyi Luo, Xintong Jiang, Yu Yang, Eugene Roy Antony Samy 외

Reliable and automated 3D plant shoot segmentation is a core prerequisite for the extraction of plant phenotypic traits at the organ level. Combining deep learning and point clouds can provide effective ways to address t…

Instance SegmentationOrgan SegmentationSegmentationSemantic Segmentation

Semantic decomposition Network with Contrastive and Structural Constraints for Dental Plaque Segmentation

2022-08-12 · Jian Shi, Baoli Sun, Xinchen Ye, Zhihui Wang 외

Segmenting dental plaque from images of medical reagent staining provides valuable information for diagnosis and the determination of follow-up treatment plan. However, accurate dental plaque segmentation is a challengin…

Segmentation

Vision-based robot manipulation of transparent liquid containers in a laboratory setting

2024-04-25 · Daniel Schober, Ronja Güldenring, James Love, Lazaros Nalpantidis

Laboratory processes involving small volumes of solutions and active ingredients are often performed manually due to challenges in automation, such as high initial costs, semi-structured environments and protocol variabi…

ArticlesRobot Manipulation

Dual Prompting for Diverse Count-level PET Denoising

2025-05-05 · Xiaofeng Liu, Yongsong Huang, Thibault Marin, Samira Vafay Eslahi 외

The to-be-denoised positron emission tomography (PET) volumes are inherent with diverse count levels, which imposes challenges for a unified model to tackle varied cases. In this work, we resort to the recently flourishe…

DenoisingPrompt Learning