paper-with-me

홈 › Papers

Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries

2017-04-12 · CVPR 2017 7 · Yuting Zhang, Luyao Yuan, Yijie Guo, Zhiyuan He, I-An Huang, Honglak Lee

Associating image regions with text queries has been recently explored as a new way to bridge visual and linguistic representations. A few pioneering approaches have been proposed based on recurrent neural language models trained generatively (e.g., generating captions), but achieving somewhat limited localization accuracy. To better address natural-language-based visual entity localization, we propose a discriminative approach. We formulate a discriminative bimodal neural network (DBNet), which can be trained by a classifier with extensive use of negative samples. Our training objective encourages better localization on single images, incorporates text phrases in a broad range, and properly pairs image regions with text phrases into positive and negative examples. Experiments on the Visual Genome dataset demonstrate the proposed DBNet significantly outperforms previous state-of-the-art methods both for localization on single images and for detection on multiple images. We we also establish an evaluation protocol for natural-language visual detection.

📄 PDF Abstract BibTeX arXiv:1704.03944

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language QueriesVisual Localization

Similar Papers 제목 키워드 기반

BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion

2021-10-04 · Zhaoqun Li, Xu Liang, Dandan Fan, Jinxing Li 외

Bimodal palmprint recognition leverages palmprint and palm vein images simultaneously,which achieves high accuracy by multi-model information fusion and has strong anti-falsification property. In the recognition pipeline…

Keypoint DetectionTranslation

Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization

2020-05-29 · Komal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan Subramanian

We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to …

Constrained Lip-synchronizationDeepFake DetectionFace SwappingTemporal Forgery Localization

Generate to Ground: Multimodal Text Conditioning Boosts Phrase Grounding in Medical Vision-Language Models

2025-07-16 · Felix Nützel, Mischa Dombrowski, Bernhard Kainz arxiv

Phrase grounding, i.e., mapping natural language phrases to specific image regions, holds significant potential for disease localization in medical imaging through clinical reports. While current state-of-the-art methods…

Phrase Grounding

Real-Time Localization and Bimodal Point Pattern Analysis of Palms Using UAV Imagery

2024-10-14 · Kangning Cui, Wei Tang, Rongkun Zhu, Manqi Wang 외

Understanding the spatial distribution of palms within tropical forests is essential for effective ecological monitoring, conservation strategies, and the sustainable integration of natural forest products into local and…

Domain-invariant NBV Planner for Active Cross-domain Self-localization

2021-02-23 · Kanji Tanaka

Pole-like landmark has received increasing attention as a domain-invariant visual cue for visual robot self-localization across domains (e.g., seasons, times of day, weathers). However, self-localization using pole-like …

Deep Reinforcement Learning