paper-with-me

홈 › Papers

VQD: Visual Query Detection in Natural Scenes

2019-04-04 · NAACL 2019 6 · Manoj Acharya, Karan Jariwala, Christopher Kanan

We propose Visual Query Detection (VQD), a new visual grounding task. In VQD, a system is guided by natural language to localize a variable number of objects in an image. VQD is related to visual referring expression recognition, where the task is to localize only one object. We describe the first dataset for VQD and we propose baseline algorithms that demonstrate the difficulty of the task compared to referring expression recognition.

📄 PDF Abstract BibTeX arXiv:1904.02794

Code (0)

등록된 구현이 없습니다.

Tasks

Referring ExpressionReferring Expression ComprehensionVisual Grounding

Similar Papers 제목 키워드 기반

Bridging the Gap between Local Semantic Concepts and Bag of Visual Words for Natural Scene Image Retrieval

2022-10-17 · Yousef Alqasrawi

This paper addresses the problem of semantic-based image retrieval of natural scenes. A typical content-based image retrieval system deals with the query image and images in the dataset as a collection of low-level featu…

Content-Based Image RetrievalImage RetrievalRetrieval

Efficient Color Boundary Detection with Color-Opponent Mechanisms

2013-06-01 · CVPR 2013 6 · Kaifu Yang, Shao-Bing Gao, Chaoyi Li, Yong-Jie Li

Color information plays an important role in better understanding of natural scenes by at least facilitating discriminating boundaries of objects or areas. In this study, we propose a new framework for boundary detection…

Boundary Detection

A DVDrive Approach for doScenes Instructed Driving Challenge

2026-06-19 · Zijian Fu, Xiangyang Chu, Mengshi Qi, Huadong Ma 외 arxiv

Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natur…

Trajectory PredictionAutonomous DrivingVisual Grounding

Zero-Query Transfer Attacks on Context-Aware Object Detectors

2022-03-29 · CVPR 2022 1 · Zikui Cai, Shantanu Rane, Alejandro E. Brito, Chengyu Song 외

Adversarial attacks perturb images such that a deep neural network produces incorrect classification results. A promising approach to defend against adversarial attacks on natural multi-object scenes is to impose a conte…

Adversarial AttackObject

VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

2025-09-10 · Chenqian Le, Yilin Zhao, Nikasadat Emami, Kushagra Yadav 외 arxiv

Recent advances in fMRI-based visual decoding have enabled compelling reconstructions of perceived images. However, most approaches rely on subject-specific training, limiting scalability and practical deployment. We int…