paper-with-me

홈 › Papers

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

2026-04-08 · Yuheng Shi, Xiaohuan Pei, Linfeng Wen, Minjing Dong, Chang Xu arxiv

MLLMs require high-resolution visual inputs for fine-grained tasks like document understanding and dense scene perception. However, current global resolution scaling paradigms indiscriminately flood the quadratic self-attention mechanism with visually redundant tokens, severely bottlenecking inference throughput while ignoring spatial sparsity and query intent. To overcome this, we propose Q-Zoom, a query-aware adaptive high-resolution perception framework that operates in an efficient coarse-to-fine manner. First, a lightweight Dynamic Gating Network safely bypasses high-resolution processing when coarse global features suffice. Second, for queries demanding fine-grained perception, a Self-Distilled Region Proposal Network (SD-RPN) precisely localizes the task-relevant Region-of-Interest (RoI) directly from intermediate feature spaces. To optimize these modules efficiently, the gating network uses a consistency-aware generation strategy to derive deterministic routing labels, while the SD-RPN employs a fully self-supervised distillation paradigm. A continuous spatio-temporal alignment scheme and targeted fine-tuning then seamlessly fuse the dense local RoI with the coarse global layout. Extensive experiments demonstrate that Q-Zoom establishes a dominant Pareto frontier. Using Qwen2.5-VL-7B as a primary testbed, Q-Zoom accelerates inference by 2.52 times on Document & OCR benchmarks and 4.39 times in High-Resolution scenarios while matching the baseline's peak accuracy. Furthermore, when configured for maximum perceptual fidelity, Q-Zoom surpasses the baseline's peak performance by 1.1% and 8.1% on these respective benchmarks. These robust improvements transfer seamlessly to Qwen3-VL, LLaVA, and emerging RL-based thinking-with-image models. Project page is available at https://yuhengsss.github.io/Q-Zoom/.

📄 PDF Abstract BibTeX arXiv:2604.06912

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

2026-06-20 · Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew, Kuan-Hao Huang 외 arxiv

Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image regions. However, we show that indiscrimina…

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

2026-02-12 · Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong 외 arxiv

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-wi…

Visual Reasoning

AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction

2026-03-05 · Jiashu Yang, Chi Zhang, Abudukelimu Wuerkaixi, Xuxin Cheng 외 arxiv

Multimodal document question answering requires retrieving dispersed evidence from visually rich long documents and performing reliable reasoning over heterogeneous information. Existing multimodal RAG systems remain lim…

Question AnsweringAnswer GenerationLogical Reasoning

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

2026-07-01 · Ziyu Ma, Shidong Yang, Yuxiang Ji, Yiming Hu 외 arxiv

High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when the image is processed as a whole. Existing methods either require tr…

Object Detection

ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks

2025-11-15 · Ruixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li 외 arxiv

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a p…

Cloud RemovalImage Editing