paper-with-me

Papers

MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments

2024-02-01 · Yang Liu, Xinshuai Song, Kaixuan Jiang, Weixing Chen, Jingzhou Luo, Guanbin Li, Liang Lin

With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically encode scene or historical memory in an unimodal manner, either visual or linguistic, which complicates the alignment of the model's action planning with embodied control. To overcome this limitation, we introduce the Multimodal Embodied Interactive Agent (MEIA), capable of translating high-level tasks expressed in natural language into a sequence of executable actions. Specifically, we propose a novel Multimodal Environment Memory (MEM) module, facilitating the integration of embodied control with large models through the visual-language memory of scenes. This capability enables MEIA to generate executable action plans based on diverse requirements and the robot's capabilities. Furthermore, we construct an embodied question answering dataset based on a dynamic virtual cafe environment with the help of the large language model. In this virtual environment, we conduct several experiments, utilizing multiple large models through zero-shot learning, and carefully design scenarios for various situations. The experimental results showcase the promising performance of our MEIA in various embodied interactive tasks.

📄 PDF Abstract BibTeX arXiv:2402.00290

Code (1)

hcplab-sysu/causalvlr 공식 구현 pytorch

Tasks

Embodied Question AnsweringLanguage ModelingLanguage ModellingLarge Language ModelQuestion AnsweringZero-Shot Learning

Similar Papers 제목 키워드 기반

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

2026-03-28 · Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia 외 arxiv

Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter do…

A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling

2026-02-04 · Dong Won Lee, Sarah Gillet, Louis-Philippe Morency, Cynthia Breazeal 외 arxiv

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minima…

A Multimodal Framework for Human-Multi-Agent Interaction

2026-03-24 · Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal arxiv

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a u…

Multimodal Reasoning

The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents

2023-09-26 · Che-Jui Chang, Samuel S. Sohn, Sen Zhang, Rajath Jayashankar 외

Previous studies regarding the perception of emotions for embodied virtual agents have shown the effectiveness of using virtual characters in conveying emotions through interactions with humans. However, creating an auto…

Large Language Models for Robotics: Opportunities, Challenges, and Perspectives

2024-01-09 · Jiaqi Wang, Zihao Wu, Yiwei Li, Hanqi Jiang 외

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and lang…

Robot Task PlanningTask Planning