paper-with-me

홈 › Papers

A Task-State Representation for Long-Horizon Mobile GUI Agents

2026-07-01 · Yujie Zheng, Zikang Liu, Xin Zhao, Ji-Rong Wen arxiv

While long-horizon mobile GUI agents typically rely on thought-action-observation loops, they struggle to separate persistent task states from transient screen observations. As execution histories grow, this entanglement imposes a severe context burden, causing agents to forget initial requirements, hallucinate progress, or repeatedly interact with stale interfaces. To address this, we introduce Task-State Representation (TSR), a training-free framework that explicitly decouples task state from sensory input. Acting as a lightweight external wrapper, TSR maintains three structured components: a global instruction summary, a dynamic progress tracker for subgoals, and a transition-aware action verifier. By continuously updating through pre- and post-action visual comparisons, TSR effectively guides the agent's reasoning without requiring architectural modifications. Experiments across four mobile GUI benchmarks validate TSR's effectiveness, yielding up to a 12 absolute point increase in success rate on complex cross-application and memory-intensive tasks.

📄 PDF Abstract BibTeX arXiv:2607.00502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoMaStage: Skill-State Graph Guided Planning and Closed-Loop Execution for Long-Horizon Indoor Mobile Manipulation

2026-03-09 · Chenxu Li, Zixuan Chen, Yetao Li, Jiapeng Xu 외 arxiv

Indoor mobile manipulation (MoMA) enables robots to translate natural language instructions into physical actions, yet long-horizon execution remains challenging due to cascading errors and limited generalization across …

MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware

2026-05-07 · Senthil Palanisamy, Abhishek Anand, Satpal Singh Rathore, Pratyush Patnaik 외 arxiv

Vision-language-action (VLA) models have driven demand for large-scale egocentric datasets, yet the hardware and infrastructure to collect long-horizon data remain inaccessible. Datasets today typically have episodes onl…

Pose Tracking

U-MASK: User-adaptive Spatio-Temporal Masking for Personalized Mobile AI Applications

2026-01-11 · Shiyuan Zhang, Yilai Liu, Yuwei Du, Ruoxuan Yang 외 arxiv

Personalized mobile artificial intelligence applications are widely deployed, yet they are expected to infer user behavior from sparse and irregular histories under a continuously evolving spatio-temporal context. This s…

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

2026-06-11 · Sunghwan Kim, Byeonghyun Pak, Kehan Long, Yulun Tian 외 arxiv

Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image observations alone. In this paper, we show …

SuperSuit: An Isomorphic Bimodal Interface for Scalable Mobile Manipulation

2026-03-06 · Tongqing Chen, Hang Wu, Jiasen Wang, Xiaotao Li 외 arxiv

High-quality, long-horizon demonstrations are essential for embodied AI, yet acquiring such data for tightly coupled wheeled mobile manipulators remains a fundamental bottleneck. Unlike fixed-base systems, mobile manipul…