paper-with-me

홈 › Papers

Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control

2025-10-16 · Zhe Wu, Hongjin Lu, Junliang Xing, Changhao Zhang, Yuxuan Li, Yin Zhu, Yuhao Yang, Yuheng Jing, Kai Li, Kun Shao, Jianye Hao, Jun Wang, Yuanchun Shi arxiv

Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on direct state-to-action mappings, which lack structured reasoning and planning, and thus generalize poorly to novel tasks or unseen UI layouts. We introduce Hi-Agent, a trainable hierarchical vision-language agent for mobile control, featuring a high-level reasoning model and a low-level action model that are jointly optimized. For efficient training, we reformulate multi-step decision-making as a sequence of single-step subgoals and propose a foresight advantage function, which leverages execution feedback from the low-level model to guide high-level optimization. This design alleviates the path explosion issue encountered by Group Relative Policy Optimization (GRPO) in long-horizon tasks and enables stable, critic-free joint training. Hi-Agent achieves a new State-Of-The-Art (SOTA) 87.9% task success rate on the Android-in-the-Wild (AitW) benchmark, significantly outperforming prior methods across three paradigms: prompt-based (AppAgent: 17.7%), supervised (Filtered BC: 54.5%), and reinforcement learning-based (DigiRL: 71.9%). It also demonstrates competitive zero-shot generalization on the ScreenSpot-v2 benchmark. On the more challenging AndroidWorld benchmark, Hi-Agent also scales effectively with larger backbones, showing strong adaptability in high-complexity mobile control scenarios.

📄 PDF Abstract BibTeX arXiv:2510.14388

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationReinforcement Learning

Similar Papers 제목 키워드 기반

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

2025-12-16 · Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato 외 arxiv

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where…

Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents

2025-10-09 · Renhua Ding, Xiao Yang, Zhengwei Fang, Jun Luo 외 arxiv

Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vulnerable to visual prompt injections, stea…

MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation

2025-07-21 · Ning Li, Xiangmou Qu, Jiamu Zhou, Jun Wang 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have enabled the development of mobile agents that can understand visual inputs and follow user instructions, unlocking new possibilities for automating complex…

MobiAgent: A Systematic Framework for Customizable Mobile Agents

2025-08-30 · Cheng Zhang, Erhu Feng, Xi Zhao, Yisheng Zhao 외 arxiv

With the rapid advancement of Vision-Language Models (VLMs), GUI-based mobile agents have emerged as a key development direction for intelligent mobile systems. However, existing agent models continue to face significant…

AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation

2025-09-29 · Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni, Jumpei Arima 외 arxiv

As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile m…