paper-with-me

홈 › Papers

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

2026-05-13 · Haoren Zhao, Tianyi Chen, Zhen Wang arxiv

Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies a critical reliability gap: state-of-the-art agents face distinct robustness challenges in real-world desktop environments characterized by multi-window stacking, occlusion, and visual clutter. To address this, we introduce WinDeskGround, a novel benchmark and synthesis framework tailored for evaluating GUI grounding robustness. Unlike static datasets, our framework parametrically generates complex desktop scenarios by controlling window occlusion, layout density, and semantic similarity, thereby simulating the distribution shifts of authentic workflows. We construct a diverse meta-dataset of 1,356 high-fidelity instruction-target pairs and conduct comprehensive evaluations of five leading MLLMs. Our results demonstrate that while top-tier agents excel in simplified settings, their accuracy declines under partial occlusion. WinDeskGround provides a valuable benchmark to facilitate the assessment and advancement of GUI agent robustness in realistic environments. The code is available at https://github.com/ZZZhr-1/WinDeskGround.

📄 PDF Abstract BibTeX arXiv:2605.16402

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

2024-02-25 · Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang 외

Despite progress in multimodal large language models (MLLMs), the challenge of interpreting long-form videos in response to linguistic queries persists, largely due to the inefficiency in temporal grounding and limited p…

Computational EfficiencyLanguage ModellingOptical Flow EstimationQuestion Answering+1

Localizing Moments in Long Video Via Multimodal Guidance

2023-02-26 · ICCV 2023 1 · Wayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron 외

The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with int…

Natural Language Moment RetrievalNatural Language Visual GroundingVideo GroundingVideo Understanding

CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding

2022-09-22 · Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao 외

This paper tackles an emerging and challenging problem of long video temporal grounding~(VTG) that localizes video moments related to a natural language (NL) query. Compared with short videos, long videos are also highly…

Contrastive LearningVideo Grounding

Generation-Guided Multi-Level Unified Network for Video Grounding

2023-03-14 · Xing Cheng, Xiangyu Wu, Dong Shen, Hezheng Lin 외

Video grounding aims to locate the timestamps best matching the query description within an untrimmed video. Prevalent methods can be divided into moment-level and clip-level frameworks. Moment-level approaches directly …

Video Grounding

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

2025-11-14 · Zihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu 외 arxiv

Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency limitations hinder practical deployment …

Object DetectionVisual Grounding