paper-with-me

홈 › Papers

Finding any Waldo: zero-shot invariant and efficient visual search

2018-07-18 · Zhang Mengmi, Feng Jiashi, Ma Keng Teck, Lim Joo Hwee, Zhao Qi, Kreiman Gabriel

Searching for a target object in a cluttered scene constitutes a fundamental challenge in daily vision. Visual search must be selective enough to discriminate the target from distractors, invariant to changes in the appearance of the target, efficient to avoid exhaustive exploration of the image, and must generalize to locate novel target objects with zero-shot training. Previous work has focused on searching for perfect matches of a target after extensive category-specific training. Here we show for the first time that humans can efficiently and invariantly search for natural objects in complex scenes. To gain insight into the mechanisms that guide visual search, we propose a biologically inspired computational model that can locate targets without exhaustive sampling and generalize to novel objects. The model provides an approximation to the mechanisms integrating bottom-up and top-down signals during search in natural scenes.

📄 PDF Abstract BibTeX arXiv:1807.10587

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

2026-05-06 · Bernhard Kainz, Johanna P Mueller, Matthew Baugh, Cosmin Bercea arxiv

Zero-shot anomaly localisation via vision-language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is fundamentally limited by the absence of healthy anatomical context. We re…

Visual Reasoning

Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions

2023-11-28 · CVPR 2024 1 · Zeyu Han, Fangrui Zhu, Qianru Lao, Huaizu Jiang

Zero-shot referring expression comprehension aims at localizing bounding boxes in an image corresponding to provided textual prompts, which requires: (i) a fine-grained disentanglement of complex visual scene and textual…

DisentanglementReferring ExpressionReferring Expression ComprehensionTriplet+1

Leveraging the Invariant Side of Generative Zero-Shot Learning

2019-04-08 · CVPR 2019 6 · Jingjing Li, Mengmeng Jin, Ke Lu, Zhengming Ding 외

Conventional zero-shot learning (ZSL) methods generally learn an embedding, e.g., visual-semantic mapping, to handle the unseen visual samples via an indirect manner. In this paper, we take the advantage of generative ad…

Generalized Zero-Shot LearningZero-Shot Learning

Efficient Zero-shot Visual Search via Target and Context-aware Transformer

2022-11-24 · Zhiwei Ding, Xuezhe Ren, Erwan David, Melissa Vo 외

Visual search is a ubiquitous challenge in natural vision, including daily tasks such as finding a friend in a crowd or searching for a car in a parking lot. Human rely heavily on relevant target features to perform goal…

To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo

2022-03-30 · Yiran Luo, Pratyay Banerjee, Tejas Gokhale, Yezhou Yang 외

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who's Waldo dataset. Given an image and a caption, PCVG requires pairing up a person's name men…

BenchmarkingPerson-centric Visual GroundingSentenceVisual Grounding