paper-with-me

Papers

Creating Multimodal Interactive Agents with Imitation and Self-Supervised Learning

2021-12-07 · DeepMind Interactive Agents Team, Josh Abramson, Arun Ahuja, Arthur Brussee, Federico Carnevale, Mary Cassin, Felix Fischer, Petko Georgiev, Alex Goldin, Mansi Gupta, Tim Harley, Felix Hill, Peter C Humphreys, Alden Hung, Jessica Landon, Timothy Lillicrap, Hamza Merzic, Alistair Muldal, Adam Santoro, Guy Scully, Tamara von Glehn, Greg Wayne, Nathaniel Wong, Chen Yan, Rui Zhu

A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through natural language. Here we study how to design artificial agents that can interact naturally with humans using the simplification of a virtual environment. We show that imitation learning of human-human interactions in a simulated world, in conjunction with self-supervised learning, is sufficient to produce a multimodal interactive agent, which we call MIA, that successfully interacts with non-adversarial humans 75% of the time. We further identify architectural and algorithmic techniques that improve performance, such as hierarchical action selection. Altogether, our results demonstrate that imitation of multi-modal, real-time human behaviour may provide a straightforward and surprisingly effective means of imbuing agents with a rich behavioural prior from which agents might then be fine-tuned for specific purposes, thus laying a foundation for training capable agents for interactive robots or digital assistants. A video of MIA's behaviour may be found at https://youtu.be/ZFgRhviF7mY

📄 PDF Abstract BibTeX arXiv:2112.03763

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningSelf-Supervised Learning

Similar Papers 제목 키워드 기반

A modular architecture for creating multimodal agents

2022-06-01 · Thomas Baier, Selene Baez Santamaria, Piek Vossen

The paper describes a flexible and modular platform to create multimodal interactive agents. The platform operates through an event-bus on which signals and interpretations are posted in a sequence in time. Different sen…

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

2025-08-24 · Aaryaman Kartha, Ahmed Masry, Mohammed Saidul Islam, Thinh Lang 외 arxiv

Dashboards are powerful visualization tools for data-driven decision-making, integrating multiple interactive views that allow users to explore, filter, and navigate data. Unlike static charts, dashboards support rich in…

Question Answering

AgentStudio: A Toolkit for Building General Virtual Agents

2024-03-26 · Longtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang 외

General virtual agents need to handle multimodal observations, master complex action spaces, and self-improve in dynamic, open-domain environments. However, existing environments are often domain-specific and require com…

Visual Grounding

Aligning to Social Norms and Values in Interactive Narratives

2022-05-04 · NAACL 2022 7 · Prithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi 외

We focus on creating agents that act in alignment with socially beneficial norms and values in interactive narratives or text-based games -- environments wherein an agent perceives and interacts with a world through natu…

text-based games

UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration

2025-09-26 · Qi Mao, Tinghan Yang, Jiahao Li, Bin Li 외 arxiv

The rapid progress of Large Multimodal Models (LMMs) and cloud-based AI agents is transforming human-AI collaboration into bidirectional, multimodal interaction. However, existing codecs remain optimized for unimodal, on…

Visual Question AnsweringText-to-Image Generation