paper-with-me

홈 › Papers

Guided Zoom: Questioning Network Evidence for Fine-grained Classification

2018-12-06 · Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang, Kate Saenko, Vittorio Murino, Stan Sclaroff

We propose Guided Zoom, an approach that utilizes spatial grounding of a model's decision to make more informed predictions. It does so by making sure the model has "the right reasons" for a prediction, defined as reasons that are coherent with those used to make similar correct decisions at training time. The reason/evidence upon which a deep convolutional neural network makes a prediction is defined to be the spatial grounding, in the pixel space, for a specific class conditional probability in the model output. Guided Zoom examines how reasonable such evidence is for each of the top-k predicted classes, rather than solely trusting the top-1 prediction. We show that Guided Zoom improves the classification accuracy of a deep convolutional neural network model and obtains state-of-the-art results on three fine-grained classification benchmark datasets.

📄 PDF Abstract BibTeX arXiv:1812.02626

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationPrediction

Similar Papers 제목 키워드 기반

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

2026-02-12 · Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong 외 arxiv

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-wi…

Visual Reasoning

TikArt: Stabilizing Aperture-Guided Fine-Grained Visual Reasoning with Reinforcement Learning

2026-02-16 · Hao Ding, Zhichuan Yang, Weijie Ge, Ziqin Gao 외 arxiv

Fine-grained visual reasoning in multimodal large language models (MLLMs) is bottlenecked by single-pass global image encoding: key evidence often lies in tiny objects, cluttered regions, subtle markings, or dense charts…

Reinforcement LearningMultimodal ReasoningVisual Reasoning

ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks

2025-11-15 · Ruixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li 외 arxiv

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a p…

Cloud RemovalImage Editing

Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in

2025-12-16 · Xiaoqian Shen, Min-Hung Chen, Yu-Chiang Frank Wang, Mohamed Elhoseiny 외 arxiv

Grounded video question answering (GVQA) aims to localize relevant temporal segments in videos and generate accurate answers to a given question; however, large video-language models (LVLMs) exhibit limited temporal awar…

Video Question AnsweringAnswer Generation

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

2026-07-28 · Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li 외 arxiv

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-releva…

Reinforcement LearningVisual Reasoning