paper-with-me

Papers

Demonstrating EMMA: Embodied MultiModal Agent for Language-guided Action Execution in 3D Simulated Environments

2022-09-01 · SIGDIAL (ACL) 2022 9 · Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, George Pantazopoulos, Amit Parekh, Arash Eshghi, Claudio Greco, Ioannis Konstas, Oliver Lemon, Verena Rieser

We demonstrate EMMA, an embodied multimodal agent which has been developed for the Alexa Prize SimBot challenge. The agent acts within a 3D simulated environment for household tasks. EMMA is a unified and multimodal generative model aimed at solving embodied tasks. In contrast to previous work, our approach treats multiple multimodal tasks as a single multimodal conditional text generation problem, where a model learns to output text given both language and visual input. Furthermore, we showcase that a single generative agent can solve tasks with visual inputs of varying length, such as answering questions about static images, or executing actions given a sequence of previous frames and dialogue utterances. The demo system will allow users to interact conversationally with EMMA in embodied dialogues in different 3D environments from the TEACh dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Text GenerationText Generation

Similar Papers 제목 키워드 기반

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

2023-11-07 · Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage 외

Interactive and embodied tasks pose at least two fundamental challenges to existing Vision & Language (VL) models, including 1) grounding language in trajectories of actions and observations, and 2) referential disambigu…

DecoderText Generation

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

2026-04-21 · Josue Torres-Fonseca, Naihao Deng, Yinpei Dai, Shane Storks 외 arxiv

Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient. We introduce SafetyALFRED, built u…

Question Answering

Visual Distraction Undermines Moral Reasoning in Vision-Language Models

2026-03-17 · Xinyi Yang, Chenheng Xu, Weijun Hong, Ce Mo 외 arxiv

Moral reasoning is fundamental to safe Artificial Intelligence (AI), yet ensuring its consistency across modalities becomes critical as AI systems evolve from text-based assistants to embodied agents. Current safety tech…

Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld

2023-11-28 · CVPR 2024 1 · Yijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao 외

While large language models (LLMs) excel in a simulated world of texts, they struggle to interact with the more realistic world without perceptions of other modalities such as visual or audio signals. Although vision-lan…

Imitation Learning

Embodied Multimodal Agents to Bridge the Understanding Gap

2021-04-01 · EACL (HCINLP) 2021 4 · Nikhil Krishnaswamy, Nada Alalyani

In this paper we argue that embodied multimodal agents, i.e., avatars, can play an important role in moving natural language processing toward “deep understanding.” Fully-featured interactive agents, model encounters bet…