paper-with-me

홈 › Papers

OmniLabel: A Challenging Benchmark for Language-Based Object Detection

2023-04-22 · ICCV 2023 1 · Samuel Schulter, Vijay Kumar B G, Yumin Suh, Konstantinos M. Dafnis, Zhixing Zhang, Shiyu Zhao, Dimitris Metaxas

Language-based object detection is a promising direction towards building a natural interface to describe objects in images that goes far beyond plain category names. While recent methods show great progress in that direction, proper evaluation is lacking. With OmniLabel, we propose a novel task definition, dataset, and evaluation metric. The task subsumes standard- and open-vocabulary detection as well as referring expressions. With more than 28K unique object descriptions on over 25K images, OmniLabel provides a challenging benchmark with diverse and complex object descriptions in a naturally open-vocabulary setting. Moreover, a key differentiation to existing benchmarks is that our object descriptions can refer to one, multiple or even no object, hence, providing negative examples in free-form text. The proposed evaluation handles the large label space and judges performance via a modified average precision metric, which we validate by evaluating strong language-based baselines. OmniLabel indeed provides a challenging test bed for future research on language-based detection.

📄 PDF Abstract BibTeX arXiv:2304.11463

Code (0)

등록된 구현이 없습니다.

Tasks

Objectobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection

2025-09-29 · Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim arxiv

Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving d…

Object Detection

Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection

2024-07-21 · KwanYong Park, Kuniaki Saito, Donghyun Kim

Vision-language (VL) models often exhibit a limited understanding of complex expressions of visual objects (e.g., attributes, shapes, and their relations), given complex and diverse language queries. Traditional approach…

Contrastive Learningobject-detectionObject DetectionSynthetic Data Generation

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

2025-03-18 · Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang 외

Large foundation models trained on large-scale vision-language data can boost Open-Vocabulary Object Detection (OVD) via synthetic training data, yet the hand-crafted pipelines often introduce bias and overfit to specifi…

DecoderObjectobject-detectionObject Detection+4

Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning

2026-03-27 · Bozhao Li, Shaocong Wu, Tong Shao, Senqiao Yang 외 arxiv

Recent advances in open-vocabulary object detection focus primarily on two aspects: scaling up datasets and leveraging contrastive learning to align language and vision modalities. However, these approaches often neglect…

Contrastive LearningObject Detection

Learning Object-Language Alignments for Open-Vocabulary Object Detection

2022-11-27 · Chuang Lin, Peize Sun, Yi Jiang, Ping Luo 외

Existing object detection methods are bounded in a fixed-set vocabulary by costly labeled data. When dealing with novel categories, the model has to be retrained with more bounding box annotations. Natural language super…

Objectobject-detectionObject DetectionOpen-vocabulary object detection+3