Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object Detection
Open-vocabulary object detection aims to detect novel categories that are independent from the base categories used during training. Most modern methods adhere to the paradigm of learning vision-language space from a large-scale multi-modal corpus and subsequently transferring the acquired knowledge to off-the-shelf detectors like Faster-RCNN. However information attenuation or destruction may occur during the process of knowledge transfer due to the domain gap hampering the generalization ability on novel categories. To mitigate this predicament in this paper we present a novel framework named BIND standing for Bulit-IN Detector to eliminate the need for module replacement or knowledge transfer to off-the-shelf detectors. Specifically we design a two-stage training framework with an Encoder-Decoder structure. In the first stage an image-text dual encoder is trained to learn region-word alignment from a corpus of image-text pairs. In the second stage a DETR-style decoder is trained to perform detection on annotated object detection datasets. In contrast to conventional manually designed non-adaptive anchors which generate numerous redundant proposals we develop an anchor proposal network that generates anchor proposals with high likelihood based on candidates adaptively thereby substantially improving detection efficiency. Experimental results on two public benchmarks COCO and LVIS demonstrate that our method stands as a state-of-the-art approach for open-vocabulary object detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Decoderobject-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object DetectionTransfer LearningWord AlignmentMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Linguistic-Aware Patch Slimming Framework for Fine-grained Cross-Modal Alignment
Cross-modal alignment aims to build a bridge connecting vision and language. It is an important multi-modal task that efficiently learns the semantic similarities between images and texts. Traditional fine-grained al…
cross-modal alignmentCross-Modal RetrievalImage RetrievalImage-to-Text Retrieval+4FSD: Fully-Specialized Detector via Neural Architecture Search
Most generic object detectors are mainly built for standard object detection tasks such as COCO and PASCAL VOC. They might not work well and/or efficiently on tasks of other domains consisting of images that are visually…
Lesion DetectionNeural Architecture SearchObjectobject-detection+1SLAN: Self-Locator Aided Network for Cross-Modal Understanding
Learning fine-grained interplay between vision and language allows to a more accurate understanding for VisionLanguage tasks. However, it remains challenging to extract key image regions according to the texts for semant…
Image RetrievalImage to textRetrievalSLAN: Self-Locator Aided Network for Vision-Language Understanding
Learning fine-grained interplay between vision and language contributes to a more accurate understanding for Vision-Language tasks. However, it remains challenging to extract key image regions according to the texts …
Image RetrievalImage to textRetrievalExploring the Vulnerability of Single Shot Module in Object Detectors via Imperceptible Background Patches
Recent works succeeded to generate adversarial perturbations on the entire image or the object of interests to corrupt CNN based object detectors. In this paper, we focus on exploring the vulnerability of the Single Shot…
ObjectRegion Proposal