VQD: Visual Query Detection in Natural Scenes
We propose Visual Query Detection (VQD), a new visual grounding task. In VQD, a system is guided by natural language to localize a variable number of objects in an image. VQD is related to visual referring expression recognition, where the task is to localize only one object. We describe the first dataset for VQD and we propose baseline algorithms that demonstrate the difficulty of the task compared to referring expression recognition.
Code (0)
등록된 구현이 없습니다.
Tasks
Referring ExpressionReferring Expression ComprehensionVisual GroundingSimilar Papers 제목 키워드 기반
Bridging the Gap between Local Semantic Concepts and Bag of Visual Words for Natural Scene Image Retrieval
This paper addresses the problem of semantic-based image retrieval of natural scenes. A typical content-based image retrieval system deals with the query image and images in the dataset as a collection of low-level featu…
Content-Based Image RetrievalImage RetrievalRetrievalEfficient Color Boundary Detection with Color-Opponent Mechanisms
Color information plays an important role in better understanding of natural scenes by at least facilitating discriminating boundaries of objects or areas. In this study, we propose a new framework for boundary detection…
Boundary DetectionA DVDrive Approach for doScenes Instructed Driving Challenge
Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natur…
Trajectory PredictionAutonomous DrivingVisual GroundingZero-Query Transfer Attacks on Context-Aware Object Detectors
Adversarial attacks perturb images such that a deep neural network produces incorrect classification results. A promising approach to defend against adversarial attacks on natural multi-object scenes is to impose a conte…
Adversarial AttackObjectVoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI
Recent advances in fMRI-based visual decoding have enabled compelling reconstructions of perceived images. However, most approaches rely on subject-specific training, limiting scalability and practical deployment. We int…