Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries
Associating image regions with text queries has been recently explored as a new way to bridge visual and linguistic representations. A few pioneering approaches have been proposed based on recurrent neural language models trained generatively (e.g., generating captions), but achieving somewhat limited localization accuracy. To better address natural-language-based visual entity localization, we propose a discriminative approach. We formulate a discriminative bimodal neural network (DBNet), which can be trained by a classifier with extensive use of negative samples. Our training objective encourages better localization on single images, incorporates text phrases in a broad range, and properly pairs image regions with text phrases into positive and negative examples. Experiments on the Visual Genome dataset demonstrate the proposed DBNet significantly outperforms previous state-of-the-art methods both for localization on single images and for detection on multiple images. We we also establish an evaluation protocol for natural-language visual detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language QueriesVisual LocalizationSimilar Papers 제목 키워드 기반
BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion
Bimodal palmprint recognition leverages palmprint and palm vein images simultaneously,which achieves high accuracy by multi-model information fusion and has strong anti-falsification property. In the recognition pipeline…
Keypoint DetectionTranslationNot made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization
We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to …
Constrained Lip-synchronizationDeepFake DetectionFace SwappingTemporal Forgery LocalizationGenerate to Ground: Multimodal Text Conditioning Boosts Phrase Grounding in Medical Vision-Language Models
Phrase grounding, i.e., mapping natural language phrases to specific image regions, holds significant potential for disease localization in medical imaging through clinical reports. While current state-of-the-art methods…
Phrase GroundingReal-Time Localization and Bimodal Point Pattern Analysis of Palms Using UAV Imagery
Understanding the spatial distribution of palms within tropical forests is essential for effective ecological monitoring, conservation strategies, and the sustainable integration of natural forest products into local and…
Domain-invariant NBV Planner for Active Cross-domain Self-localization
Pole-like landmark has received increasing attention as a domain-invariant visual cue for visual robot self-localization across domains (e.g., seasons, times of day, weathers). However, self-localization using pole-like …
Deep Reinforcement Learning