paper-with-me

Papers

An Interactive Agent Foundation Model

2024-02-08 · Zane Durante, Bidipta Sarkar, Ran Gong, Rohan Taori, Yusuke Noda, Paul Tang, Ehsan Adeli, Shrinidhi Kowshika Lakshmikanth, Kevin Schulman, Arnold Milstein, Demetri Terzopoulos, Ade Famoti, Noboru Kuno, Ashley Llorens, Hoi Vo, Katsu Ikeuchi, Li Fei-Fei, Jianfeng Gao, Naoki Wake, Qiuyuan Huang

The development of artificial intelligence systems is transitioning from creating static, task-specific models to dynamic, agent-based systems capable of performing well in a wide range of applications. We propose an Interactive Agent Foundation Model that uses a novel multi-task agent training paradigm for training AI agents across a wide range of domains, datasets, and tasks. Our training paradigm unifies diverse pre-training strategies, including visual masked auto-encoders, language modeling, and next-action prediction, enabling a versatile and adaptable AI framework. We demonstrate the performance of our framework across three separate domains -- Robotics, Gaming AI, and Healthcare. Our model demonstrates its ability to generate meaningful and contextually relevant outputs in each area. The strength of our approach lies in its generality, leveraging a variety of data sources such as robotics sequences, gameplay data, large-scale video datasets, and textual information for effective multimodal and multi-task learning. Our approach provides a promising avenue for developing generalist, action-taking, multimodal systems.

📄 PDF Abstract BibTeX arXiv:2402.05929

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingmodelMulti-Task Learning

Similar Papers 제목 키워드 기반

UItron: Foundational GUI Agent with Advanced Perception and Planning

2025-08-29 · Zhixiong Zeng, Jing Huang, Liming Zheng, Wenkang Han 외 arxiv

GUI agent aims to enable automated operations on Mobile/PC devices, which is an important task toward achieving artificial general intelligence. The rapid advancement of VLMs accelerates the development of GUI agents, ow…

Reinforcement Learning

Evaluating Cognitive Age Alignment in Interactive AI Agents

2026-05-18 · Yifan Shen, Jiawen Zhang, Jian Xu, Junho Kim 외 arxiv

While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across domains ranging from daily life to advanced scientific research, a profo…

Visual Reasoning

Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities

2026-02-04 · Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li 외 arxiv

Uncertainty quantification (UQ) for large language models (LLMs) is a key building block for safety guardrails of daily LLM applications. Yet, even as LLM agents are increasingly deployed in highly complex tasks, most UQ…

Neural Interactive Proofs

2024-12-12 · Lewis Hammond, Sam Adam-Day

We consider the problem of how a trusted, but computationally bounded agent (a 'verifier') can learn to interact with one or more powerful but untrusted agents ('provers') in order to solve a given task. More specificall…

Position: Foundation Agents as the Paradigm Shift for Decision Making

2024-05-27 · Xiaoqian Liu, Xingzhou Lou, Jianbin Jiao, Junge Zhang

Decision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches to decision making face challenges related to low sample efficiency and poor gene…

Decision MakingPosition