paper-with-me

홈 › Papers

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

2025-05-16 · Bin Lei, Weitai Kang, Zijian Zhang, Winson Chen, Xi Xie, Shan Zuo, Mimi Xie, Ali Payani, Mingyi Hong, Yan Yan, Caiwen Ding

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricate workflows around a single large model or only provide workflow modularity, our agent integrates tool-based and pure vision agents within a highly modular architecture, enabling different models to collaboratively solve decoupled tasks in a step-by-step manner. Our generality is demonstrated by our ability to evaluate not only pure vision-based real-world benchmarks (i.e., OSWorld), but also more general or tool-intensive benchmarks (e.g., GAIA and SWE-Bench). Specifically, we achieve $\mathbf{7.27\%}$ accuracy on OSWorld, higher than Claude-Computer-Use. Codes and evaluation scripts are open-sourced at https://github.com/bin123apple/InfantAgent.

📄 PDF Abstract BibTeX arXiv:2505.10887

Code (1)

bin123apple/infantagent 공식 구현

Similar Papers 제목 키워드 기반

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

2025-11-11 · Zhengyang Liang, Daoan Zhang, Huichi Zhou, Rui Huang 외 arxiv

While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilities. To bridge this gap, we introduce Un…

Object Segmentation

Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task Experts

2025-06-12 · Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 외

Recently, agents based on multimodal large language models (MLLMs) have achieved remarkable progress across various domains. However, building a generalist agent with capabilities such as perception, planning, action, gr…

DiversityMinecraftMixture-of-ExpertsMultimodal Reasoning

Unlocking Generalization for Robotics via Modularity and Scale

2025-03-10 · Murtaza Dalal

How can we build generalist robot systems? Scale may not be enough due to the significant multimodality of robotics tasks, lack of easily accessible data and the challenges of deploying on physical hardware. Meanwhile, m…

Scene Generation

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

2024-12-11 · CVPR 2025 1 · Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev 외

We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus …

An Interactive Agent Foundation Model

2024-02-08 · Zane Durante, Bidipta Sarkar, Ran Gong, Rohan Taori 외

The development of artificial intelligence systems is transitioning from creating static, task-specific models to dynamic, agent-based systems capable of performing well in a wide range of applications. We propose an Int…

Language ModelingLanguage ModellingmodelMulti-Task Learning