paper-with-me

홈 › Papers

AttZoom: Attention Zoom for Better Visual Features

2025-08-05 · Daniel DeAlcala, Aythami Morales, Julian Fierrez, Ruben Tolosana arxiv

We present Attention Zoom, a modular and model-agnostic spatial attention mechanism designed to improve feature extraction in convolutional neural networks (CNNs). Unlike traditional attention approaches that require architecture-specific integration, our method introduces a standalone layer that spatially emphasizes high-importance regions in the input. We evaluated Attention Zoom on multiple CNN backbones using CIFAR-100 and TinyImageNet, showing consistent improvements in Top-1 and Top-5 classification accuracy. Visual analyses using Grad-CAM and spatial warping reveal that our method encourages fine-grained and diverse attention patterns. Our results confirm the effectiveness and generality of the proposed layer for improving CCNs with minimal architectural overhead.

📄 PDF Abstract BibTeX arXiv:2508.03625

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

2026-06-20 · Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew, Kuan-Hao Huang 외 arxiv

Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image regions. However, we show that indiscrimina…

Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements

2026-03-15 · Ziwei Liu, Tao Feng, Borui Kang, Yanbing Yang 외 arxiv

Multimodal Large Language Model (MLLM)-based Graphical User Interface (GUI) agents develop rapidly, with visual grounding that maps natural language instructions to target UI elements serving as the core capability. Exis…

Visual Grounding

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

2026-07-28 · Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li 외 arxiv

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-releva…

Reinforcement LearningVisual Reasoning

Localizing Anatomical Landmarks in Ocular Images using Zoom-In Attentive Networks

2022-09-25 · Xiaofeng Lei, Shaohua Li, Xinxing Xu, Huazhu Fu 외

Localizing anatomical landmarks are important tasks in medical image analysis. However, the landmarks to be localized often lack prominent visual features. Their locations are elusive and easily confused with the backgro…

Medical Image Analysisobject-detectionObject Detection

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling

2025-09-28 · Xianjie Liu, Yiman Hu, Yixiong Zou, Liang Wu 외 arxiv

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding tasks. However, their performance on high-resolution images remains suboptimal. While existing approaches often attribute thi…