paper-with-me

홈 › Papers

A Survey on (M)LLM-Based GUI Agents

2025-03-27 · Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, Yueting Zhuang

Graphical User Interface (GUI) Agents have emerged as a transformative paradigm in human-computer interaction, evolving from rule-based automation scripts to sophisticated AI-driven systems capable of understanding and executing complex interface operations. This survey provides a comprehensive examination of the rapidly advancing field of LLM-based GUI Agents, systematically analyzing their architectural foundations, technical components, and evaluation methodologies. We identify and analyze four fundamental components that constitute modern GUI Agents: (1) perception systems that integrate text-based parsing with multimodal understanding for comprehensive interface comprehension; (2) exploration mechanisms that construct and maintain knowledge bases through internal modeling, historical experience, and external information retrieval; (3) planning frameworks that leverage advanced reasoning methodologies for task decomposition and execution; and (4) interaction systems that manage action generation with robust safety controls. Through rigorous analysis of these components, we reveal how recent advances in large language models and multimodal learning have revolutionized GUI automation across desktop, mobile, and web platforms. We critically examine current evaluation frameworks, highlighting methodological limitations in existing benchmarks while proposing directions for standardization. This survey also identifies key technical challenges, including accurate element localization, effective knowledge retrieval, long-horizon planning, and safety-aware execution control, while outlining promising research directions for enhancing GUI Agents' capabilities. Our systematic review provides researchers and practitioners with a thorough understanding of the field's current state and offers insights into future developments in intelligent interface automation.

📄 PDF Abstract BibTeX arXiv:2504.13865

Code (0)

등록된 구현이 없습니다.

Tasks

Action GenerationInformation RetrievalRetrievalSurvey

Similar Papers 제목 키워드 기반

From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes

2026-04-24 · Rubén Garzón, Pauline Baron, Vincent Grari, Jonne Kamphorst 외 arxiv

Large language models (LLM) agents may offer tools to predict human responses to surveys. A common technique for defining these agents uses only demographics, for example country, age, gender, employment status, income, …

SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation

2026-02-11 · Beichen Guo, Zhiyuan Wen, Jia Gu, Haochen Shi 외 arxiv

Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialized ASG frameworks and Deep Research agent…

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

2024-12-05 · Xiachong Feng, Longxu Dou, Ella Li, Qinghao Wang 외

Game-theoretic scenarios have become pivotal in evaluating the social intelligence of Large Language Model (LLM)-based social agents. While numerous studies have explored these agents in such settings, there is a lack of…

Language ModelingLanguage ModellingLarge Language ModelSurvey

Large Language Model-Based Agents for Software Engineering: A Survey

2024-09-04 · Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng 외

The recent advance in Large Language Models (LLMs) has shaped a new paradigm of AI agents, i.e., LLM-based agents. Compared to standalone LLMs, LLM-based agents substantially extend the versatility and expertise of LLMs …

AI AgentLanguage ModelingLanguage ModellingLarge Language Model+1

A Survey on Large Language Model based Autonomous Agents

2023-08-22 · Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang 외

Autonomous agents have long been a prominent research focus in both academic and industry communities. Previous research in this field often focuses on training agents with limited knowledge within isolated environments,…

Language ModelingLanguage ModellingLarge Language ModelSurvey