paper-with-me

Papers

Learning a Semantic Calibration Network for Open-Vocabulary Semantic Segmentation

2026-06-06 · Yang Sun, Tao Wang, Anastasia Ioannou, Ge Xu arxiv

Semantic image segmentation assigns a predefined category label to each pixel, has achieved significant progress lately. Open-Vocabulary Segmentation (OVS) extends the segmentation task from a fixed set to an open set, enabling the identification and segmentation of novel concepts based on arbitrary text inputs, such as category names or descriptions. In this paper, we propose a novel Semantic Calibration Network (SCN) for open-vocabulary semantic segmentation. Different from prior approaches that focus on feature aggregation or simple fine-tuning of pre-trained models, SCN refines the mask classification process by explicitly modeling the semantic correlations between classes, aiming to enhance the model's discriminative power while effectively preserving the generalization abilities of the pre-trained CLIP model. Specifically, SCN comprises two core components: Class Disambiguation (CD) and Logits Fusion (LF). First, a cross-attention mechanism is utilized to transform the text embeddings into visually aware pseudo-text embeddings, in order to derive an enhanced similarity score that complements the original mask-text similarity score. Subsequently, the Class Disambiguation module captures implicit inter-class dependencies through a residual architecture to effectively resolve semantic ambiguities. Finally, the Logits Fusion module dynamically integrates multifaceted semantic evidence to ensure that the model achieves a robust semantic consensus while maintaining CLIP's inherent generalization capability. Comprehensive experimental results on mainstream benchmarks demonstrate that the proposed method achieves significant performance improvements compared to state-of-the-art algorithms.

📄 PDF Abstract BibTeX arXiv:2606.08001

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationImage Segmentation

Similar Papers 제목 키워드 기반

Open-Vocabulary Segmentation with Semantic-Assisted Calibration

2023-12-07 · CVPR 2024 1 · Yong liu, Sule Bai, Guanbin Li, Yitong Wang 외

This paper studies open-vocabulary segmentation (OVS) through calibrating in-vocabulary and domain-biased embedding space with generalized contextual prior of CLIP. As the core of open-vocabulary understanding, alignment…

AttributeOpen Vocabulary Semantic Segmentation

Seeking Consensus: Geometric-Semantic On-the-Fly Recalibration for Open-Vocabulary Remote Sensing Semantic Segmentation

2026-04-29 · Guanchun Wang, Chenxiao Wu, Xiangrong Zhang, Zelin Peng 외 arxiv

Open-vocabulary semantic segmentation (OVSS) in remote sensing images is a promising task that employs textual descriptions for identifying undefined land cover categories. Despite notable advances, existing methods typi…

Semantic Segmentation

Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

2026-07-02 · Hu Zhu, Bohan Li, Xianda Guo, Yanlun Peng 외 arxiv

Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without relying on external camera calibration. Rece…

Novel View SynthesisScene Understanding

DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF

2025-07-19 · Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot 외 arxiv

3D semantic segmentation provides high-level scene understanding for applications in robotics, autonomous systems, \textit{etc}. Traditional methods adapt exclusively to either task-specific goals (open-vocabulary segmen…

Unsupervised Semantic Segmentation3D Semantic SegmentationScene Understanding

Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization

2026-07-01 · Xuying Huang, Sicong Pan, Maren Bennewitz arxiv

Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains underexplored in open-vocabulary 3D semantic segmentation. Existing m…

3D Semantic SegmentationScene Understanding