paper-with-me

홈 › Papers

dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3

2026-03-19 · Saikat Dutta, Biplab Banerjee, Hamid Rezatofighi arxiv

Open-Vocabulary Semantic Segmentation (OVSS) assigns pixel-level labels from an open set of text-defined categories, demanding reliable generalization to unseen classes at inference. Although modern vision-language models (VLMs) support strong open-vocabulary recognition, their representations learned through global contrastive objectives remain suboptimal for dense prediction, prompting many OVSS methods to depend on limited adaptation or refinement of image-text similarity maps. This, in turn, restricts spatial precision and robustness in complex, cluttered scenes. We introduce dinov3.seg, extending dinov3.txt into a dedicated framework for OVSS. Our contributions are four-fold. First, we design a task-specific architecture tailored to this backbone, systematically adapting established design principles from prior open-vocabulary segmentation work. Second, we jointly leverage text embeddings aligned with both the global [CLS] token and local patch-level visual features from ViT-based encoder, effectively combining semantic discrimination with fine-grained spatial locality. Third, unlike prior approaches that rely primarily on post hoc similarity refinement, we perform early refinement of visual representations prior to image-text interaction, followed by late refinement of the resulting image-text correlation features, enabling more accurate and robust dense predictions in cluttered scenes. Finally, we propose a high-resolution local-global inference strategy based on sliding-window aggregation, which preserves spatial detail while maintaining global context. We conduct extensive experiments on five widely adopted OVSS benchmarks to evaluate our approach. The results demonstrate its effectiveness and robustness, consistently outperforming current state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2603.19531

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation

Similar Papers 제목 키워드 기반

DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery

2026-05-04 · Ryan Faulkenberry, Saurabh Prasad arxiv

The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery well without supervised fine-tuning are valuable, but existing solution…

Open Vocabulary Semantic SegmentationFeature Upsampling

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

2024-11-28 · Luca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara 외

Open-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leve…

Segmentation

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

2024-12-20 · CVPR 2025 1 · Cijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre 외

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP, self-supervised visual fe…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentationzero-shot-classification+1

LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes

2024-10-18 · Juliette Marrie, Romain Menegaux, Michael Arbel, Diane Larlus 외

We address the problem of extending the capabilities of vision foundation models such as DINO, SAM, and CLIP, to 3D tasks. Specifically, we introduce a novel method to uplift 2D image features into Gaussian Splatting rep…

3D geometryobject-detectionObject DetectionObject Localization+2

Streamlined Open-Vocabulary Human-Object Interaction Detection

2026-03-29 · Chang Sun, Dongliang Liao, Changxing Ding arxiv

Open-vocabulary human-object interaction (HOI) detection aims to localize and recognize all human-object interactions in an image, including those unseen during training. Existing approaches usually rely on the collabora…

Human-Object Interaction Detection