paper-with-me

홈 › Papers

Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection

2019-07-23 · ICCV 2019 10 · Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li, Danfeng Qin, Jesse Berent

Learning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the need for boxes to that of image-level annotations, even cheaper supervision is naturally available in the form of unstructured textual descriptions that users may freely provide when uploading image content. However, straightforward approaches to using such data for WSOD wastefully discard captions that do not exactly match object names. Instead, we show how to squeeze the most information out of these captions by training a text-only classifier that generalizes beyond dataset boundaries. Our discovery provides an opportunity for learning detection models from noisy but more abundant and freely-available caption data. We also validate our model on three classic object detection benchmarks and achieve state-of-the-art WSOD performance. Our code is available at https://github.com/yekeren/Cap2Det.

📄 PDF Abstract BibTeX arXiv:1907.10164

Code (1)

yekeren/Cap2Det 공식 구현 tf

Tasks

Objectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Learning Better Visual Representations for Weakly-Supervised Object Detection Using Natural Language Supervision

2021-09-29 · Mesut Erhan Unal, Adriana Kovashka

We present a framework to better leverage natural language supervision for a specific downstream task, namely weakly-supervised object detection (WSOD). Our framework employs a multimodal pre-training step, during which …

cross-modal alignmentobject-detectionObject DetectionRepresentation Learning+1

Multi-source weak supervision for saliency detection

2019-04-01 · CVPR 2019 6 · Yu Zeng, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang 외

The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-…

Caption GenerationSaliency DetectionSaliency Prediction

Learning Object Detection from Captions via Textual Scene Attributes

2020-09-30 · Achiya Jerbi, Roei Herzig, Jonathan Berant, Gal Chechik 외

Object detection is a fundamental task in computer vision, requiring large annotated datasets that are difficult to collect, as annotators need to label objects and their bounding boxes. Thus, it is a significant challen…

Image CaptioningObjectobject-detectionObject Detection

Learning to discover and localize visual objects with open vocabulary

2018-11-25 · Keren Ye, Mingda Zhang, Wei Li, Danfeng Qin 외

To alleviate the cost of obtaining accurate bounding boxes for training today's state-of-the-art object detection models, recent weakly supervised detection work has proposed techniques to learn from image-level labels. …

Objectobject-detectionObject Detection

Open-Vocabulary Object Detection Using Captions

2020-11-20 · CVPR 2021 1 · Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, Shih-Fu Chang

Despite the remarkable accuracy of deep neural networks in object detection, they are costly to train and scale due to supervision requirements. Particularly, learning more object categories typically requires proportion…

Objectobject-detectionObject DetectionOpen Vocabulary Attribute Detection+3