paper-with-me

홈 › Papers

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception

2025-11-25 · Miguel Carvalho, Helder Dias, Bruno Martins arxiv

Vision-Language Models (VLMs) often struggle with tasks that require fine-grained image understanding, such as scene-text recognition or document analysis, due to perception limitations and visual fragmentation. To address these challenges, we introduce CropVLM as an external low-cost method for boosting performance, enabling VLMs to dynamically ''zoom in'' on relevant image regions, enhancing their ability to capture fine details. CropVLM is trained using reinforcement learning, without using human-labeled bounding boxes as a supervision signal, and without expensive synthetic evaluations. The model is trained once and can be paired with both open-source and proprietary VLMs to improve their performance. Our approach delivers significant improvements on tasks that require high-resolution image understanding, notably for benchmarks that are out-of-domain for the target VLM, without modifying or fine-tuning the VLM, thus avoiding catastrophic forgetting.

📄 PDF Abstract BibTeX arXiv:2511.19820

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis

2026-05-05 · Abderrahmene Boudiaf, Sajd Javed arxiv

High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck," where manual data collection is labor-int…

Zero-shot Generalization

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

2026-02-12 · Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong 외 arxiv

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-wi…

Visual Reasoning

Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding

2025-12-05 · Zhiyuan Jiang, Shenghao Xie, Wenyi Li, Wenqiang Zu 외 arxiv

Grounding is a fundamental capability for building graphical user interface (GUI) agents. Although existing approaches rely on large-scale bounding box supervision, they still face various challenges, such as cross-platf…

ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks

2025-11-15 · Ruixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li 외 arxiv

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a p…

Cloud RemovalImage Editing

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

2024-11-25 · Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu 외

An image, especially with high-resolution, typically consists of numerous visual elements, ranging from dominant large objects to fine-grained detailed objects. When perceiving such images, multimodal large language mode…

AI AgentVisual Question AnsweringVisual Question Answering (VQA)