paper-with-me

홈 › Papers

Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs

2026-03-14 · Nimrod Shabtay, Moshe Kimhi, Artem Spector, Sivan Haray, Ehud Rivlin, Chaim Baskin, Raja Giryes, Eli Schwartz arxiv

Vision-language models (VLMs) typically process images at a native high-resolution, forcing a trade-off between accuracy and computational efficiency: high-resolution inputs capture fine details but incur significant computational costs, while low-resolution inputs advocate for efficiency, they potentially miss critical visual information, like small text. We present AwaRes, a spatial-on-demand framework that resolves this accuracy-efficiency trade-off by operating on a low-resolution global view and using tool-calling to retrieve only high-resolution segments needed for a given query. We construct supervised data automatically: a judge compares low- vs.\ high-resolution answers to label whether cropping is needed, and an oracle grounding model localizes the evidence for the correct answer, which we map to a discrete crop set to form multi-turn tool-use trajectories. We train our framework with cold-start SFT followed by multi-turn GRPO with a composite reward that combines semantic answer correctness with explicit crop-cost penalties. Project page: https://nimrodshabtay.github.io/AwaRes

📄 PDF Abstract BibTeX arXiv:2603.16932

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search

2025-11-25 · Yunqi Zhou, Chengjie Jiang, Chun Yuan, Jing Li arxiv

With advances in satellite constellations, sensor technologies, and imaging pipelines, ultra-high-resolution (Ultra-HR) remote sensing imagery is becoming increasingly widespread. However, current remote sensing foundati…

Visual Question Answering

HRDA: Context-Aware High-Resolution Domain-Adaptive Semantic Segmentation

2022-04-27 · Lukas Hoyer, Dengxin Dai, Luc van Gool

Unsupervised domain adaptation (UDA) aims to adapt a model trained on the source domain (e.g. synthetic data) to the target domain (e.g. real-world data) without requiring further annotations on the target domain. This w…

Domain AdaptationGPUImage-to-Image TranslationSegmentation+4

Small data deep learning methodology for in-field disease detection

2024-09-25 · David Herrera-Poyato, Jacinto Domínguez-Rull, Rosana Montes, Inés Hernánde 외

Early detection of diseases in crops is essential to prevent harvest losses and improve the quality of the final product. In this context, the combination of machine learning and proximity sensors is emerging as a techni…

Data AugmentationDeep Learning

Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens

2026-01-09 · Yohann Perron, Vladyslav Sydorov, Christophe Pottier, Loic Landrieu arxiv

Current approaches for segmenting ultra high resolution images either slide a window, thereby discarding global context, or downsample and lose fine detail. We propose a simple yet effective method that brings explicit m…

Semantic Segmentation

Annual field-scale maps of tall and short crops at the global scale using GEDI and Sentinel-2

2022-12-19 · Stefania Di Tommaso, Sherrie Wang, Vivek Vajipey, Noel Gorelick 외

Crop type maps are critical for tracking agricultural land use and estimating crop production. Remote sensing has proven an efficient and reliable tool for creating these maps in regions with abundant ground labels for m…