paper-with-me

홈 › Papers

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning

2026-05-08 · Yubin Wu, Zicheng Cai, Liping Ning, Hua Wang, Zhi Chen, Yaohua Tang, Hao Chen arxiv

Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device agents are constrained by limited model capacity, and further performance improvements remain urgently needed. Traditional Supervised Fine-Tuning (SFT) for small-scale models often leads to overfitting, catastrophic forgetting and policy rigidity, and thus fails to fully address these challenges. In this work, we propose a novel SFT-free training paradigm that significantly enhances the performance of small-scale models. We first present the initial systematic integration of generalized knowledge distillation into the GUI agent domain via Guided On-policy Distillation. By incorporating oracle reference trajectories together with a dynamic retrieval mechanism, our method reduces hallucinations and mitigates the cognitive misalignment inherent in multi-solution GUI tasks. Building on this foundation, we further introduce a Multi-solution Dual-level GRPO framework that jointly aligns macro-level subtask planning with micro-level execution matching, thereby improving exploration in long-horizon GUI agent scenarios. In addition, we construct an automated data generation pipeline to synthesize GUI task trajectories with rich multi-solution annotations. Extensive experiments show that our method achieves state-of-the-art performance among lightweight models while remaining competitive with substantially larger-scale models across all benchmarks. Ablation studies further demonstrate that structured on-policy distillation and multi-solution dual-level exploration can fully unlock the capabilities of 2B/3B scale agents, surpassing the performance limits of conventional imitation learning.

📄 PDF Abstract BibTeX arXiv:2605.07505

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningKnowledge Distillation

Similar Papers 제목 키워드 기반

Information-Bottleneck-Based Behavior Representation Learning for Multi-agent Reinforcement learning

2021-09-29 · Yue Jin, Shuangqing Wei, Jian Yuan, Xudong Zhang

In multi-agent deep reinforcement learning, extracting sufficient and compact information of other agents is critical to attain efficient convergence and scalability of an algorithm. In canonical frameworks, distilling o…

Deep Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+2

Distilling LLMs' Decomposition Abilities into Compact Language Models

2024-02-02 · Denis Tarasov, Kumar Shridhar

Large Language Models (LLMs) have demonstrated proficiency in their reasoning abilities, yet their large size presents scalability challenges and limits any further customization. In contrast, compact models offer custom…

TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents

2025-07-02 · Dmytro Kuzmenko, Nadiya Shvai arxiv

We present a novel approach to knowledge transfer in model-based reinforcement learning, addressing the critical challenge of deploying large world models in resource-constrained environments. Our method efficiently dist…

Reinforcement Learning

Distilling Deep RL Models Into Interpretable Neuro-Fuzzy Systems

2022-09-07 · Arne Gevaert, Jonathan Peck, Yvan Saeys

Deep Reinforcement Learning uses a deep neural network to encode a policy, which achieves very good performance in a wide range of applications but is widely regarded as a black box model. A more interpretable alternativ…

Deep Reinforcement LearningOpenAI Gymreinforcement-learningReinforcement Learning+1

MapDream: Task-Driven Map Learning for Vision-Language Navigation

2026-01-30 · Guoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang 외 arxiv

Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception. Howev…

Vision-Language Navigation