paper-with-me

홈 › Papers

MintAct: A Unified Visual Agent for Digital Environments

2026-09-18 · Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege Özsoy, Kaixin Ma, Vishwesh Kirthivasan, Oğuzhan Fatih Kar, Roman Bachmann, Anders Boesen Lindbo Larsen, Afshin Dehghan hf

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.

📄 PDF Abstract BibTeX arXiv:2609.22083

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

KG-MAS: Knowledge Graph-Enhanced Multi-Agent Infrastructure for coupling physical and digital robotic environments

2025-10-11 · Walid Abdela arxiv

The seamless integration of physical and digital environments in Cyber-Physical Systems(CPS), particularly within Industry 4.0, presents significant challenges stemming from system heterogeneity and complexity. Tradition…

Visually-grounded Humanoid Agents

2026-04-09 · Hang Ye, Xiaoxuan Ma, Fan Lu, Wayne Wu 외 arxiv

Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which li…

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

2024-12-13 · Zhiqi Ge, Juncheng Li, Xinglei Pang, Minghe Gao 외

Digital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based agents built on Large Language Models (LLMs…

Edge Detection

Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

2025-06-18 · Yining Hong, Rui Sun, Bingxuan Li, Xingcheng Yao 외

AI agents today are mostly siloed - they either retrieve and reason over vast amount of digital information and knowledge obtained online; or interact with the physical world through embodied perception, planning and act…

CocoaBench: Evaluating Unified Digital Agents in the Wild

2026-04-13 · CocoaBench Team, Shibo Hao, Zhining Zhang, Zhiqi Liang 외 arxiv

LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly integrating these capabilities into unified…

Visual Grounding