paper-with-me

홈 › Papers

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

2026-06-19 · Vincent Siu, Manasi Sharma, Dawn Song, Daniel Yue Zhang, Chenguang Wang arxiv

Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap with ChainWorld, which composes atomic OSWorld tasks into long horizon desktop workloads through directional compatibility search while preserving the source evaluators. The resulting workload contains 347 chains of length two to four and compares two renderings of the same task sequence. In single turn evaluation, all tasks are presented together in one prompt. In multi turn evaluation, tasks are revealed one at a time. Across four current computer use agents, maximum chain completion is 31%. Multi turn evaluation improves completion for three models, but both protocols remain challenging. The two protocols also expose different failure profiles. Single turn failures concentrate on artifact precision, while multi turn failures more often reflect session management problems such as fragmented progress and later turn disengagement.

📄 PDF Abstract BibTeX arXiv:2606.21654

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

2026-06-02 · Wenkai Wang, Tao Xiong, Jingchen Ni, Yunpeng Bao 외 arxiv

Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information a…

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

2026-06-08 · Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu 외 arxiv

Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these inte…

ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents

2026-06-22 · Yincheng Zhou, Athena Zhuoming Zhong, Shijie Zhang, Kevin Zhang 외 arxiv

As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes a key challenge. GUI tasks require long-horizon sequences of precise m…

IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use Agents

2026-02-19 · Seoyoung Lee, Seobin Yoon, Seongbeen Lee, Yoojung Chun 외 arxiv

Computer-use agents operate over long horizons under noisy perception, multi-window contexts, evolving environment states. Existing approaches, from RL-based planners to trajectory retrieval, often drift from user intent…

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

2026-07-21 · Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu 외 hf

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet v…