paper-with-me

Papers

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

2026-02-06 · Yichen Gong, Zhuohan Cai, Sunhao Dai, Yuqi Zhou, Zhangxuan Gu, Changhua Meng, Shuheng Shen arxiv

Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of real-world mobile usage. To this end, we introduce VenusBench-Mobile, a challenging online benchmark for evaluating general-purpose mobile GUI agents under realistic, user-centric conditions. VenusBench-Mobile builds two core evaluation pillars: defining what to evaluate via user-intent-driven task design that reflects real mobile usage, and how to evaluate through a capability-oriented annotation scheme for fine-grained agent behavior analysis. Extensive evaluation of state-of-the-art mobile GUI agents reveals large performance gaps relative to prior benchmarks, indicating that VenusBench-Mobile poses substantially more challenging and realistic tasks and that current agents remain far from reliable real-world deployment. Diagnostic analysis further shows that failures are dominated by deficiencies in perception and memory, which are largely obscured by coarse-grained evaluations. Moreover, even the strongest agents exhibit near-zero success under environment variations, highlighting their brittleness in realistic settings. Based on these insights, we believe VenusBench-Mobile provides an important stepping stone toward robust real-world deployment of mobile GUI agents. Code and data are available at https://github.com/inclusionAI/UI-Venus/tree/VenusBench-Mobile.

📄 PDF Abstract BibTeX arXiv:2604.06182

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UI-Venus-1.5 Technical Report

2026-02-09 · Venus Team, Changlong Gao, Zhangxuan Gu, Yulin Liu 외 arxiv

GUI agents have emerged as a powerful paradigm for automating interactions in digital environments, yet achieving both broad generality and consistently strong task performance remains challenging. In this report, we pre…

Reinforcement Learning

VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks

2025-12-18 · Beitong Zhou, Zhexiao Huang, Yuan Guo, Zhangxuan Gu 외 arxiv

GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, o…

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

2026-08-24 · Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu 외 arxiv

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps,…

Reinforcement Learning

MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware

2026-05-07 · Senthil Palanisamy, Abhishek Anand, Satpal Singh Rathore, Pratyush Patnaik 외 arxiv

Vision-language-action (VLA) models have driven demand for large-scale egocentric datasets, yet the hardware and infrastructure to collect long-horizon data remain inaccessible. Datasets today typically have episodes onl…

Pose Tracking

Towards Decentralized Task Offloading and Resource Allocation in User-Centric Mobile Edge Computing

2023-12-03 · Langtian Qin, Hancheng Lu, Yuang Chen, Baolin Chong 외

In the traditional cellular-based mobile edge computing (MEC), users at the edge of the cell are prone to suffer severe inter-cell interference and signal attenuation, leading to low throughput even transmission interrup…

Deep Reinforcement LearningEdge-computing