paper-with-me

홈 › Papers

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

2025-10-31 · Tao Liu, Chongyu Wang, Rongjie Li, Yingchen Yu, Xuming He, Bai Song arxiv

While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction, and history summarization. The structured reasoning component generates coherent Chain-of-Thought analyses combining progress estimation and decision reasoning, which inform both immediate action predictions and compact history summaries for future steps. Based on this framework, we train a GUI agent, \textbf{GUI-Rise}, through supervised fine-tuning on pseudo-labeled trajectories and reinforcement learning with Group Relative Policy Optimization (GRPO). This framework employs specialized rewards, including a history-aware objective, directly linking summary quality to subsequent action performance. Comprehensive evaluations on standard benchmarks demonstrate state-of-the-art results under identical training data conditions, with particularly strong performance in out-of-domain scenarios. These findings validate our framework's ability to maintain robust reasoning and generalization across diverse GUI navigation tasks. Code is available at https://leon022.github.io/GUI-Rise.

📄 PDF Abstract BibTeX arXiv:2510.27210

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDomain Generalization

Similar Papers 제목 키워드 기반

BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation

2026-06-19 · Rithvik Jonna, Aakash Gurram, Man Namgung, Wyatt Mackey 외 arxiv

Vision-Language Models (VLMs) for embodied navigation rely on selecting a fixed number of frames from a growing trajectory history. As episodes extend, this selection grows increasingly sparse, yet prior work shows no ac…

Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning

2025-11-12 · Mobin Habibpour, Fatemeh Afghah arxiv

While Vision-Language Models (VLMs) are set to transform robotic navigation, existing methods often underutilize their reasoning capabilities. To unlock the full potential of VLMs in robotics, we shift their role from pa…

History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

2025-12-16 · Xichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang 외 arxiv

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both …

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

2025-06-19 · Mobin Habibpour, Fatemeh Afghah

Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ t…

AgentEHR: Advancing Autonomous Clinical Decision-Making via Retrospective Summarization

2026-01-20 · Yusheng Liao, Chuan Xuan, Yutong Cai, Lina Yang 외 arxiv

Large Language Models have demonstrated profound utility in the medical domain. However, their application to autonomous Electronic Health Records~(EHRs) navigation remains constrained by a reliance on curated inputs and…