paper-with-me

홈 › Papers

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

2026-01-12 · Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Qingyun Li, Yian Wang, Yu Qiao, Zun Wang, Zichen Ding arxiv

While Vision-Language Models (VLMs) have significantly advanced Computer-Using Agents (CUAs), current frameworks struggle with robustness in long-horizon workflows and generalization in novel domains. These limitations stem from a lack of granular control over historical visual context curation and the absence of visual-aware tutorial retrieval. To bridge these gaps, we introduce OS-Symphony, a holistic framework that comprises an Orchestrator coordinating two key innovations for robust automation: (1) a Reflection-Memory Agent that utilizes milestone-driven long-term memory to enable trajectory-level self-correction, effectively mitigating visual context loss in long-horizon tasks; (2) Versatile Tool Agents featuring a Multimodal Searcher that adopts a SeeAct paradigm to navigate a browser-based sandbox to synthesize live, visually aligned tutorials, thereby resolving fidelity issues in unseen scenarios. Experimental results demonstrate that OS-Symphony delivers substantial performance gains across varying model scales, establishing new state-of-the-art results on three online benchmarks, notably achieving 65.84% on OSWorld.

📄 PDF Abstract BibTeX arXiv:2601.07779

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly

2026-01-30 · Wei Zhu, Zhiwen Tang, Kun Yue arxiv

Recent advancements have increasingly focused on leveraging large language models (LLMs) to construct autonomous agents for complex problem-solving tasks. However, existing approaches predominantly employ a single-agent …

Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence

2025-08-27 · Ji Wang, Kashing Chen, Xinyuan Song, Ke Zhang 외 arxiv

Most existing Large Language Model (LLM)-based agent frameworks rely on centralized orchestration, incurring high deployment costs, rigid communication topologies, and limited adaptability. To address these challenges, w…

Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures

2025-10-18 · Minh-Khoi Nguyen-Nhat, Rachel S. Y. Teo, Laziz Abdullaev, Maurice Mok 외 arxiv

Sparse Mixture of Experts (SMoE) has emerged as a promising solution to achieving unparalleled scalability in deep learning by decoupling model parameter count from computational cost. By activating only a small subset o…

Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding

2026-03-18 · Haiyang Yan, Hongyun Zhou, Peng Xu, Xiaoxue Feng 외 arxiv

Despite rapid developments and widespread applications of MLLM agents, they still struggle with long-form video understanding (LVU) tasks, which are characterized by high information density and extended temporal spans. …

Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

2025-04-01 · Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang 외

Computer use agents automate digital tasks by directly interacting with graphical user interfaces (GUIs) on computers and mobile devices, offering significant potential to enhance human productivity by completing an open…

AI AgentTask Planning