paper-with-me

홈 › Papers

Revisiting Visual Grounding

2019-04-03 · WS 2019 6 · Erik Conser, Kennedy Hahn, Chandler M. Watson, Melanie Mitchell

We revisit a particular visual grounding method: the "Image Retrieval Using Scene Graphs" (IRSG) system of Johnson et al. (2015). Our experiments indicate that the system does not effectively use its learned object-relationship models. We also look closely at the IRSG dataset, as well as the widely used Visual Relationship Dataset (VRD) that is adapted from it. We find that these datasets exhibit biases that allow methods that ignore relationships to perform relatively well. We also describe several other problems with the IRSG dataset, and report on experiments using a subset of the dataset in which the biases and other problems are removed. Our studies contribute to a more general effort: that of better understanding what machine learning methods that combine language and vision actually learn and what popular datasets actually test.

📄 PDF Abstract BibTeX arXiv:1904.02225

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrievalVisual Grounding

Similar Papers 제목 키워드 기반

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning

2026-06-25 · Mengzhao Wang, Yanli Ji, Wangmeng Zuo, Peng Ye 외 arxiv

Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeatedly forwarding selected visual tokens. A …

Multimodal ReasoningVisual Grounding

Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization

2025-11-27 · Yifan Du, Kun Zhou, Yingqian Min, Yue Ling 외 arxiv

We study how different Chain-of-Thought (CoT) designs affect the acquisition of the generalizable visual reasoning ability in vision-language models (VLMs). While CoT data, especially long or visual CoT such as "think wi…

Visual Reasoning

Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning

2026-02-18 · Mingjia Shi, Yinhan He, Yaochen Zhu, Jundong Li arxiv

Vision-language models (VLMs) aim to reason by jointly leveraging visual and textual modalities. While allocating additional inference-time computation has proven effective for large language models (LLMs), achieving sim…

Visual Grounding

Improving Visual Reasoning with Iterative Evidence Refinement

2026-03-14 · Zeru Shi, Kai Mei, Yihao Quan, Dimitris N. Metaxas 외 arxiv

Vision language models (VLMs) are increasingly capable of reasoning over images, but robust visual reasoning often requires re-grounding intermediate steps in the underlying visual evidence. Recent approaches typically r…

Reinforcement LearningVisual Reasoning

Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields

2025-10-03 · Zhiting Mei, Ola Shorinwa, Anirudha Majumdar arxiv

Semantic distillation in radiance fields has spurred significant advances in open-vocabulary robot policies, e.g., in manipulation and navigation, founded on pretrained semantics from large vision models. While prior wor…

Object LocalizationPose Estimation