paper-with-me

홈 › Papers

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

2026-04-06 · Weihao Cao, Runqi Wang, Xiaoyue Duan, Jinchao Zhang, Ang Yang, Liping Jing arxiv

Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods achieve strong detection performance on general scenarios (e.g., OV-COCO) but suffer severe performance drops when transferred to downstream tasks with substantial domain shifts. This degradation stems from the scarcity and weak semantics of category labels in domain-specific task, as well as the inability of existing models to capture auxiliary semantics beyond coarse-grained category label. To address these issues, we propose HSA-DINO, a parameter-efficient semantic augmentation framework for enhancing open-vocabulary object detection. Specifically, we propose a multi-scale prompt bank that leverages image feature pyramids to capture hierarchical semantics and select domain-specific local semantic prompts, progressively enriching textual representations from coarse to fine-grained levels. Furthermore, we introduce a semantic-aware router that dynamically selects the appropriate semantic augmentation strategy during inference, thereby preventing parameter updates from degrading the generalization ability of the pre-trained OVOD model. We evaluate HSA-DINO on OV-COCO, several vertical domain datasets, and modified benchmark settings. The results show that HSA-DINO performs favorably against previous state-of-the-art methods, achieving a superior trade-off between domain adaptability and open-vocabulary generalization.

📄 PDF Abstract BibTeX arXiv:2604.04444

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

LOSC: LiDAR Open-voc Segmentation Consolidator

2025-07-10 · Nermin Samet, Gilles Puy, Renaud Marlet arxiv

We study the use of image-based Vision-Language Models (VLMs) for open-vocabulary segmentation of lidar scans in driving settings. Classically, image semantics can be back-projected onto 3D point clouds. Yet, resulting p…

Panoptic SegmentationPoint Clouds

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

2025-04-28 · CVPR 2025 1 · Yan Wang, Baoxiong Jia, Ziyu Zhu, Siyuan Huang

Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a nov…

3D Semantic SegmentationContrastive LearningScene UnderstandingSemantic Segmentation

O2V-Mapping: Online Open-Vocabulary Mapping with Neural Implicit Representation

2024-04-10 · Muer Tie, Julong Wei, Zhengjun Wang, Ke wu 외

Online construction of open-ended language scenes is crucial for robotic applications, where open-vocabulary interactive scene understanding is required. Recently, neural implicit representation has provided a promising …

Image SegmentationObjectObject LocalizationScene Understanding+2

Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing

2026-03-02 · Zijin Yin, Bing Li, Kongming Liang, Hao Sun 외 arxiv

Semantic segmentation takes pivotal roles in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is critical to test their behaviors in varied an…

Semantic SegmentationAutonomous DrivingData AugmentationStyle Transfer

Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation

2025-09-01 · Jiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang 외 arxiv

Open-vocabulary semantic segmentation (OVSS) conducts pixel-level classification via text-driven alignment, where the domain discrepancy between base category training and open-vocabulary inference poses challenges in di…

Semantic Segmentation