paper-with-me

홈 › Papers

PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

2026-03-18 · Jianjian Yin, Tao Chen, Yi Chen, Gensheng Pei, Xiangbo Shu, Yazhou Yao, Fumin Shen arxiv

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial structure of spatial and class aggregations, leading to knowledge interference between class-level semantics and spatial context. Therefore, this paper proposes a simple yet effective parallel cost aggregation (PCA-Seg) paradigm to alleviate the above challenge, enabling the model to capture richer vision-language alignment information from cost volumes. Specifically, we design an expert-driven perceptual learning (EPL) module that efficiently integrates semantic and contextual streams. It incorporates a multi-expert parser to extract complementary features from multiple perspectives. In addition, a coefficient mapper is designed to adaptively learn pixel-specific weights for each feature, enabling the integration of complementary knowledge into a unified and robust feature embedding. Furthermore, we propose a feature orthogonalization decoupling (FOD) strategy to mitigate redundancy between the semantic and contextual streams, which allows the EPL module to learn diverse knowledge from orthogonalized features. Extensive experiments on eight benchmarks show that each parallel block in PCA-Seg adds merely 0.35M parameters while achieving state-of-the-art OSPS performance.

📄 PDF Abstract BibTeX arXiv:2603.17520

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

2023-03-21 · CVPR 2024 1 · Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab 외

Open-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions. In this work, we introduce a novel cost-based approach to adapt vision-langu…

Image SegmentationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentation+2

DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery

2026-05-04 · Ryan Faulkenberry, Saurabh Prasad arxiv

The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery well without supervised fine-tuning are valuable, but existing solution…

Open Vocabulary Semantic SegmentationFeature Upsampling

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

2025-09-15 · Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao 외 arxiv

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluat…

Dimensionality ReductionImage SegmentationDomain Adaptation

Effective SAM Combination for Open-Vocabulary Semantic Segmentation

2024-11-22 · CVPR 2025 1 · Minhyeok Lee, Suhwan Cho, Jungho Lee, Sunghun Yang 외

Open-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such…

DecoderLanguage ModelingLanguage ModellingOpen Vocabulary Semantic Segmentation+3

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

2026-01-04 · Shengji Tang, Weihao Lin, Peng Ye, Jingqi Ye 외 arxiv

Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that op…