paper-with-me

홈 › Papers

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding

2026-05-04 · Ruilin Yao, Shegnwu Xiong, Tianyu Zou, Shili Xiong, Yi Rong arxiv

Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. However, grounding performance degrades in high-resolution interfaces, where dense layouts and small interactive elements expose a resolution gap between modern displays and model input constraints. Existing zoom-in strategies rely on fixed anchors, heuristic grids, or reinforcement learning, lacking a principled mechanism to adaptively determine where refinement is needed and how much spatial uncertainty should be explored. We propose AutoFocus, a training-free, uncertainty-aware active visual search framework for GUI grounding. Our key insight is that token-level perplexity in coordinate generation naturally reflects spatial uncertainty. Rather than committing to a single prediction, AutoFocus samples multiple coordinate hypotheses and converts their axial perplexities into an anisotropic gaussian spatial probability field, explicitly modeling directional uncertainty. Based on this field, we generate global and local region proposals and introduce Shape-Aware Zooming to balance tight localization with contextual preservation. A visual prompt-based aggregation step then selects the most consistent prediction via structured comparison. Extensive experiments on ScreenSpot-Pro and ScreenSpot-V2 demonstrate consistent improvements across both general-purpose and GUI-specialized VLMs.

📄 PDF Abstract BibTeX arXiv:2605.02630

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations

2025-11-23 · Litian Gong, Fatemeh Bahrani, Yutai Zhou, Amin Banayeeanzade 외 arxiv

AutoFocus-IL is a simple yet effective method to improve data efficiency and generalization in visual imitation learning by guiding policies to attend to task-relevant features rather than distractors and spurious correl…

Robot Manipulation

An End-to-End Autofocus Camera for Iris on the Move

2021-06-29 · Leyuan Wang, Kunbo Zhang, Yunlong Wang, Zhenan Sun

For distant iris recognition, a long focal length lens is generally used to ensure the resolution ofiris images, which reduces the depth of field and leads to potential defocus blur. To accommodate users at different dis…

Iris Recognition

Intensity-Robust Autofocus for Spike Camera

2024-01-01 · CVPR 2024 1 · Changqing Su, Zhiyuan Ye, Yongsheng Xiao, You Zhou 외

Spike cameras a novel neuromorphic visual sensor can capture full-time spatial information through spike stream offering ultra-high temporal resolution and an extensive dynamic range. Autofocus control (AC) plays a p…

Synthetic Defocus and Look-Ahead Autofocus for Casual Videography

2019-05-15 · Xuaner Zhang, Kevin Matzen, Vivien Nguyen, Dillon Yao 외

In cinema, large camera lenses create beautiful shallow depth of field (DOF), but make focusing difficult and expensive. Accurate cinema focus usually relies on a script and a person to control focus in realtime. Casual …

BIG-bench Machine LearningSaliency Detection

Versatile optimization-based speed-up method for autofocusing in digital holographic microscopy

2023-05-17 · Julianna Winnik, Damian Suski, Piotr Zdańkowski, Luiza Stanaszek 외

We propose a speed-up method for the in-focus plane detection in digital holographic microscopy that can be applied to a broad class of autofocusing algorithms that involve repetitive propagation of an object wave to var…

Position