paper-with-me

홈 › Papers

LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training

2025-10-16 · Yiming Wang, Da Yin, Yuedong Cui, Ruichen Zheng, Zhiqian Li, Zongyu Lin, Di Wu, Xueqing Wu, Chenchen Ye, Yu Zhou, Kai-Wei Chang arxiv

Digital agents require diverse, large-scale UI trajectories to generalize across real-world tasks, yet collecting such data is prohibitively expensive in both human annotation, infra and engineering perspectives. To this end, we introduce $\textbf{UI-Simulator}$, a scalable paradigm that generates structured UI states and transitions to synthesize training trajectories at scale. Our paradigm integrates a digital world simulator for diverse UI states, a guided rollout process for coherent exploration, and a trajectory wrapper that produces high-quality and diverse trajectories for agent training. We further propose $\textbf{UI-Simulator-Grow}$, a targeted scaling strategy that enables more rapid and data-efficient scaling by prioritizing high-impact tasks and synthesizes informative trajectory variants. Experiments on WebArena and AndroidWorld show that UI-Simulator rivals or surpasses open-source agents trained on real UIs with significantly better robustness, despite using weaker teacher models. Moreover, UI-Simulator-Grow matches the performance of Llama-3-70B-Instruct using only Llama-3-8B-Instruct as the base model, highlighting the potential of targeted synthesis scaling paradigm to continuously and efficiently enhance the digital agents.

📄 PDF Abstract BibTeX arXiv:2510.14969

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis

2025-12-18 · Zhi Helu, Huang Jingjing, Xu Wang, Xu Yangbin 외 arxiv

Embodied intelligence, a grand challenge in artificial intelligence, is fundamentally constrained by the limited spatial understanding and reasoning capabilities of current models. Prevailing efforts to address this thro…

Spatial Reasoning

LuciBot: Automated Robot Policy Learning from Generated Videos

2025-03-12 · Xiaowen Qiu, Yian Wang, Jiting Cai, Zhehuan Chen 외

Automatically generating training supervision for embodied tasks is crucial, as manual designing is tedious and not scalable. While prior works use large language models (LLMs) or vision-language models (VLMs) to generat…

Video Generation

CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators

2026-05-09 · Nicolás Astorga, Anita Kriz, Mihaela van der Schaar arxiv

Despite surpassing human performance across mathematics, coding, and other knowledge-intensive tasks, large language models (LLMs) continue to struggle with causal reasoning. A core obstacle is the target data itself: ca…

Data Augmentation

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

2024-02-18 · Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei 외

This paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs), aiming for a more accurate assessment of their capabilities and limitations. We utilize a mu…

Model Selection

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

2026-03-30 · Muyang He, Hanzhong Guo, Junxiong Lin, Yizhou Yu arxiv

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between…

Autonomous DrivingVideo Generation