Read, look and detect: Bounding box annotation from image-caption pairs
Various methods have been proposed to detect objects while reducing the cost of data annotation. For instance, weakly supervised object detection (WSOD) methods rely only on image-level annotations during training. Unfortunately, data annotation remains expensive since annotators must provide the categories describing the content of each image and labeling is restricted to a fixed set of categories. In this paper, we propose a method to locate and label objects in an image by using a form of weaker supervision: image-caption pairs. By leveraging recent advances in vision-language (VL) models and self-supervised vision transformers (ViTs), our method is able to perform phrase grounding and object detection in a weakly supervised manner. Our experiments demonstrate the effectiveness of our approach by achieving a 47.51% recall@1 score in phrase grounding on Flickr30k Entities and establishing a new state-of-the-art in object detection by achieving 21.1 mAP 50 and 10.5 mAP 50:95 on MS COCO when exclusively relying on image-caption pairs.
Code (0)
등록된 구현이 없습니다.
Tasks
Objectobject-detectionObject DetectionPhrase GroundingWeakly Supervised Object DetectionSimilar Papers 제목 키워드 기반
Associative embeddings for large-scale knowledge transfer with self-assessment
We propose a method for knowledge transfer between semantically related classes in ImageNet. By transferring knowledge from the images that have bounding-box annotations to the others, our method is capable of automatica…
Gaussian ProcessesObject LocalizationTransfer LearningClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection
High-quality annotations are essential for object detection models, but ensuring label accuracy - especially for bounding boxes - remains both challenging and costly. This paper introduces ClipGrader, a novel approach th…
Objectobject-detectionObject DetectionPseudo Label+1Iterative Bounding Box Annotation for Object Detection
Manual annotation of bounding boxes for object detection in digital images is tedious, and time and resource consuming. In this paper, we propose a semi-automatic method for efficient bounding box annotation. The method …
Objectobject-detectionObject DetectionSegmentation-Based Bounding Box Generation for Omnidirectional Pedestrian Detection
We propose a segmentation-based bounding box generation method for omnidirectional pedestrian detection that enables detectors to tightly fit bounding boxes to pedestrians without omnidirectional images for training. Due…
object-detectionObject DetectionPedestrian DetectionTraining object class detectors with click supervision
Training object class detectors typically requires a large set of images with objects annotated by bounding boxes. However, manually drawing bounding boxes is very time consuming. In this paper we greatly reduce annotati…
Multiple Instance LearningObjectObject LocalizationWeakly-Supervised Object Localization