paper-with-me

Papers

GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior

2025-06-09 · Penghao Wu, Shengnan Ma, Bo wang, Jiaheng Yu, Lewei Lu, Ziwei Liu

Multimodal Large Language Models (MLLMs) have shown great potential in revolutionizing Graphical User Interface (GUI) automation. However, existing GUI models mostly rely on learning from nearly error-free offline trajectories, thus lacking reflection and error recovery capabilities. To bridge this gap, we propose GUI-Reflection, a novel framework that explicitly integrates self-reflection and error correction capabilities into end-to-end multimodal GUI models throughout dedicated training stages: GUI-specific pre-training, offline supervised fine-tuning (SFT), and online reflection tuning. GUI-reflection enables self-reflection behavior emergence with fully automated data generation and learning processes without requiring any human annotation. Specifically, 1) we first propose scalable data pipelines to automatically construct reflection and error correction data from existing successful trajectories. While existing GUI models mainly focus on grounding and UI understanding ability, we propose the GUI-Reflection Task Suite to learn and evaluate reflection-oriented abilities explicitly. 2) Furthermore, we built a diverse and efficient environment for online training and data collection of GUI models on mobile devices. 3) We also present an iterative online reflection tuning algorithm leveraging the proposed environment, enabling the model to continuously enhance its reflection and error correction abilities. Our framework equips GUI agents with self-reflection and correction capabilities, paving the way for more robust, adaptable, and intelligent GUI automation, with all data, models, environments, and tools to be released publicly.

📄 PDF Abstract BibTeX arXiv:2506.08012

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

2025-06-02 · Zhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang 외

Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle with complex problems requiring explicit self-reflection and self-correction, especially compared to their…

Multimodal Reasoningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

ReBeCA: Unveiling Interpretable Behavior Hierarchy behind the Iterative Self-Reflection of Language Models with Causal Analysis

2026-02-06 · Tianqiang Yan, Sihan Shang, Yuheng Li, Song Qiu 외 arxiv

While self-reflection can enhance language model reliability, its underlying mechanisms remain opaque, with existing analyses often yielding correlation-based insights that fail to generalize. To address this, we introdu…

ReflCtrl: Controlling LLM Reflection via Representation Engineering

2025-12-16 · Ge Yan, Chung-En Sun, Tsui-Wei, Weng arxiv

Large language models (LLMs) with Chain-of-Thought (CoT) reasoning have achieved strong performance across diverse tasks, including mathematics, coding, and general reasoning. A distinctive ability of these reasoning mod…

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models

2025-06-13 · XUDONG ZHU, Jiachen Jiang, Mohammad Mahdi Khalili, Zhihui Zhu

Self-reflection -- the ability of a large language model (LLM) to revisit, evaluate, and revise its own reasoning -- has recently emerged as a powerful behavior enabled by reinforcement learning with verifiable rewards (…

Large Language ModelNavigate

From Latent Signals to Reflection Behavior: Tracing Meta-Cognitive Activation Trajectory in R1-Style LLMs

2026-02-02 · Yanrui Du, Yibo Gao, Sendong Zhao, Jiayun Li 외 arxiv

R1-style LLMs have attracted growing attention for their capacity for self-reflection, yet the internal mechanisms underlying such behavior remain unclear. To bridge this gap, we anchor on the onset of reflection behavio…