paper-with-me

Papers

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

2024-12-16 · Qi Sun, Pengfei Hong, Tej Deep Pala, Vernon Toh, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to generate actionable policies tailored to specific robotic embodiments. To address this, Visual-Language-Action (VLA) models have emerged, yet they face challenges in long-horizon spatial reasoning and grounded task planning. In this work, we propose the Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning, Emma-X. Emma-X leverages our constructed hierarchical embodiment dataset based on BridgeV2, containing 60,000 robot manipulation trajectories auto-annotated with grounded task reasoning and spatial guidance. Additionally, we introduce a trajectory segmentation strategy based on gripper states and motion trajectories, which can help mitigate hallucination in grounding subtask reasoning generation. Experimental results demonstrate that Emma-X achieves superior performance over competitive baselines, particularly in real-world robotic tasks requiring spatial reasoning.

📄 PDF Abstract BibTeX arXiv:2412.11974

Code (1)

declare-lab/emma-x 공식 구현 pytorch

Tasks

HallucinationRobot ManipulationScene UnderstandingSpatial ReasoningTask Planning

Similar Papers 제목 키워드 기반

Demonstrating EMMA: Embodied MultiModal Agent for Language-guided Action Execution in 3D Simulated Environments

2022-09-01 · SIGDIAL (ACL) 2022 9 · Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, George Pantazopoulos 외

We demonstrate EMMA, an embodied multimodal agent which has been developed for the Alexa Prize SimBot challenge. The agent acts within a 3D simulated environment for household tasks. EMMA is a unified and multimodal gene…

Conditional Text GenerationText Generation

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

2023-11-07 · Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage 외

Interactive and embodied tasks pose at least two fundamental challenges to existing Vision & Language (VL) models, including 1) grounding language in trajectories of actions and observations, and 2) referential disambigu…

DecoderText Generation

A Multimodal Framework for Human-Multi-Agent Interaction

2026-03-24 · Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal arxiv

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a u…

Multimodal Reasoning

Embodied Multimodal Agents to Bridge the Understanding Gap

2021-04-01 · EACL (HCINLP) 2021 4 · Nikhil Krishnaswamy, Nada Alalyani

In this paper we argue that embodied multimodal agents, i.e., avatars, can play an important role in moving natural language processing toward “deep understanding.” Fully-featured interactive agents, model encounters bet…

Agent AI: Surveying the Horizons of Multimodal Interaction

2024-01-07 · Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong 외

Multi-modal AI systems will likely become a ubiquitous presence in our everyday lives. A promising approach to making these systems more interactive is to embody them as agents within physical and virtual environments. A…

multimodal interaction