paper-with-me

홈 › Papers

Enhance Mobile Agents Thinking Process Via Iterative Preference Learning

2025-05-18 · Kun Huang, Weikai Xu, Yuxuan Liu, Quandong Wang, Pengzhi Gao, Wei Liu, Jian Luan, Bin Wang, Bo An

The Chain of Action-Planning Thoughts (CoaT) paradigm has been shown to improve the reasoning performance of VLM-based mobile agents in GUI tasks. However, the scarcity of diverse CoaT trajectories limits the expressiveness and generalization ability of such agents. While self-training is commonly employed to address data scarcity, existing approaches either overlook the correctness of intermediate reasoning steps or depend on expensive process-level annotations to construct process reward models (PRM). To address the above problems, we propose an Iterative Preference Learning (IPL) that constructs a CoaT-tree through interative sampling, scores leaf nodes using rule-based reward, and backpropagates feedback to derive Thinking-level Direct Preference Optimization (T-DPO) pairs. To prevent overfitting during warm-up supervised fine-tuning, we further introduce a three-stage instruction evolution, which leverages GPT-4o to generate diverse Q\&A pairs based on real mobile UI screenshots, enhancing both generality and layout understanding. Experiments on three standard Mobile GUI-agent benchmarks demonstrate that our agent MobileIPL outperforms strong baselines, including continual pretraining models such as OS-ATLAS and UI-TARS. It achieves state-of-the-art performance across three standard Mobile GUI-Agents benchmarks and shows strong generalization to out-of-domain scenarios.

📄 PDF Abstract BibTeX arXiv:2505.12299

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Pretraining

Methods 이 논문이 사용한 방법론

+ ( 1 ) ⟷ 888 ⟷ ( 829 ) ⟷ 0881||How do I resolve a dispute on Expedia? How do I resolve a dispute on Expedia contact their support at + ( 1 ) ⟷ 888 ⟷ ( 829 ) ⟷ 0881 or + ( 1 ) ⟷ 805 ⟷ ( 330 ) ⟷ 4056. Provide booking details and explain the issue…
CoaT Co-Scale Conv-Attentional Image Transformer (CoaT) is a Transformer-based image classifier equipped with co-scale and…

Similar Papers 제목 키워드 기반

ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

2025-03-12 · Ziyu Wan, Yunxiang Li, Xiaoyu Wen, Yan Song 외

Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking -- enabling models to monitor, evaluate, and control their reasoning processes for…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning

Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards

2025-06-25 · Jihao Gu, Qihang Ai, Yingyao Wang, Pi Bu 외

Vision-language model-based mobile agents have gained the ability to not only understand complex instructions and mobile screenshots, but also optimize their action outputs via thinking and reasoning, benefiting from rei…

reinforcement-learningReinforcement Learning

Iterative Forward Tuning Boosts In-Context Learning in Language Models

2023-05-22 · Jiaxi Yang, Binyuan Hui, Min Yang, Bailin Wang 외

Despite the advancements in in-context learning (ICL) for large language models (LLMs), current research centers on specific prompt engineering, such as demonstration selection, with the expectation that a single iterati…

Decision MakingIn-Context LearningMultiple-choicePrompt Engineering

MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism

2025-11-14 · Shulin Liu, Dong Du, Tao Yang, Yang Li 외 arxiv

Recent progress in large language models (LLMs) has been propelled by reinforcement learning with verifiable rewards (RLVR) and test-time scaling. However, the limited output length of LLMs constrains the depth of reason…

Reinforcement Learning

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

2024-11-04 · Biao Wu, Yanda Li, Meng Fang, Zirui Song 외

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and process multimodal data have grown. This su…

multimodal interactionSurvey