paper-with-me

홈 › Papers

GUI-PRA: Process Reward Agent for GUI Tasks

2025-09-27 · Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang arxiv

Graphical User Interface (GUI) Agents powered by Multimodal Large Language Models (MLLMs) show significant potential for automating tasks. However, they often struggle with long-horizon tasks, leading to frequent failures. Process Reward Models (PRMs) are a promising solution, as they can guide these agents with crucial process signals during inference. Nevertheless, their application to the GUI domain presents unique challenges. When processing dense artificial inputs with long history data, PRMs suffer from a "lost in the middle" phenomenon, where the overwhelming historical context compromises the evaluation of the current step. Furthermore, standard PRMs lacks GUI changing awareness, providing static evaluations that are disconnected from the dynamic consequences of actions, a critical mismatch with the inherently dynamic nature of GUI tasks. In response to these challenges, we introduce GUI-PRA (Process Reward Agent for GUI Tasks), a judge agent designed to better provide process reward than standard PRM by intelligently processing historical context and actively perceiving UI state changes. Specifically, to directly combat the ``lost in the middle'' phenomenon, we introduce a dynamic memory mechanism consisting of two core components: a Relevance-based Retrieval Module to actively fetch pertinent information from long histories and a Progressive Summarization Module to dynamically condense growing interaction data, ensuring the model focuses on relevant context. Moreover, to address the lack of UI changing awareness, we introduce an Aadaptive UI Perception mechanism. This mechanism enables the agent to reason about UI state changes and dynamically select the most appropriate tool to gather grounded visual evidence, ensuring its evaluation is always informed by the current UI context.

📄 PDF Abstract BibTeX arXiv:2509.23263

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RF-Agent: Automated Reward Function Design via Language Agent Tree Search

2026-02-27 · Ning Gao, Xiuhui Zhang, Xingyu Jiang, Mukang You 외 arxiv

Designing efficient reward functions for low-level control tasks is a challenging problem. Recent research aims to reduce reliance on expert experience by using Large Language Models (LLMs) with task information to gener…

GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks

2025-09-28 · Cong Chen, Kaixiang Ji, Hao Zhong, Muzhi Zhu 외 arxiv

Autonomous agents for long-sequence Graphical User Interface tasks are hindered by sparse rewards and the intractable credit assignment problem. To address these challenges, we introduce GUI-Shepherd, a Process Reward Mo…

Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning

2024-12-19 · Aditya Kapoor, Sushant Swamy, Kale-ab Tessera, Mayank Baranwal 외

In multi-agent environments, agents often struggle to learn optimal policies due to sparse or delayed global rewards, particularly in long-horizon tasks where it is challenging to evaluate actions at intermediate time st…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningTAR

AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories

2025-04-11 · Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel 외

Web agents enable users to perform tasks on web browsers through natural language interaction. Evaluating web agents trajectories is an important problem, since it helps us determine whether the agent successfully comple…

Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

2024-06-17 · Weimin Xiong, YiFan Song, Xiutian Zhao, Wenhao Wu 외

Large language model agents have exhibited exceptional performance across a range of complex interactive tasks. Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they prima…

Language ModelingLanguage ModellingLarge Language Model