paper-with-me

홈 › Papers

FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting

2026-04-30 · Fengxian Ji, Jingpu Yang, Zirui Song, Yuanxi Wang, Zhexuan Cui, Yuke Li, Qian Jiang, Xiuying Chen arxiv

Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target-state definitions, and an overreliance on final-task success, obscuring where and why agents fail. To address this gap, we introduce \textbf{FineState-Bench}, a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. FineState-Bench comprises 2,209 instances across desktop, web, and mobile platforms, spanning four interaction families and 23 UI component types, with each instance explicitly specifying an exact target state for fine-grained state setting. We further propose \textit{FineState-Metrics}, a four-stage diagnostic pipeline with stage-wise success rates: Localization Success Rate (SR@Loc), Interaction Success Rate (SR@Int), Exact State Success Rate at Locate (ES-SR@Loc), and Exact State Success Rate at Interact (ES-SR@Int), and a plug-and-play \textit{Visual Diagnostic Assistant} (VDA) that generates a Description and a bounding-box Localization Hint to diagnose visual grounding reason via controlled w/ vs.\ w/o comparisons. On FineState-Bench, exact goal-state success remains low: ES-SR@Int peaks at 32.8\% on Web and 22.8\% on average across platforms. With VDA localization hints, Gemini-2.5-Flash gains +14.9 ES-SR@Int points, suggesting substantial headroom from improved visual grounding, yet overall accuracy is still insufficient for reliable fine-grained state-conditioned interaction \href{https://github.com/FengxianJi/FineState-Bench}{Github.}

📄 PDF Abstract BibTeX arXiv:2604.27974

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents

2025-08-12 · Fengxian Ji, Jingpu Yang, Zirui Song, Yuanxi Wang 외 arxiv

With the rapid advancement of generative artificial intelligence technology, Graphical User Interface (GUI) agents have demonstrated tremendous potential for autonomously managing daily tasks through natural language ins…

Visual Localization

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

2026-05-31 · Jingjie Lin, Bingbing Wang, Zihan Wang, Zhengda Jin 외 arxiv

Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure the reflective memory required to synthesize fragmented, multimodal cue…

VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement

2026-01-20 · Tiancheng Fang, Bowen Pan, Lingxi Chen, Jiangjing Lyu 외 arxiv

We propose VIAFormer, a Voxel-Image Alignment Transformer model designed for Multi-view Conditioned Voxel Refinement--the task of repairing incomplete noisy voxels using calibrated multi-view images as guidance. Its effe…

How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos

2026-04-16 · Olga Loginova, Frank Keller arxiv

Reliable procedural monitoring in video requires exposure to naturally occurring human errors and the recoveries that follow. In egocentric recordings, mistakes are often partially occluded by hands and revealed through …

Video Generation

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

2026-06-02 · Litao Liu, Yifan Han, Pengfei Yi, Wenbo Yu 외 arxiv

Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dependent and often one-to-many in cluttered scenes: the same object ma…