paper-with-me

Papers

Multimodal Contextualized Plan Prediction for Embodied Task Completion

2023-05-10 · Mert İnan, Aishwarya Padmakumar, Spandana Gella, Patrick Lange, Dilek Hakkani-Tur

Task planning is an important component of traditional robotics systems enabling robots to compose fine grained skills to perform more complex tasks. Recent work building systems for translating natural language to executable actions for task completion in simulated embodied agents is focused on directly predicting low level action sequences that would be expected to be directly executable by a physical robot. In this work, we instead focus on predicting a higher level plan representation for one such embodied task completion dataset - TEACh, under the assumption that techniques for high-level plan prediction from natural language are expected to be more transferable to physical robot systems. We demonstrate that better plans can be predicted using multimodal context, and that plan prediction and plan execution modules are likely dependent on each other and hence it may not be ideal to fully decouple them. Further, we benchmark execution of oracle plans to quantify the scope for improvement in plan prediction models.

📄 PDF Abstract BibTeX arXiv:2305.06485

Code (0)

등록된 구현이 없습니다.

Tasks

PredictionTask Planning

Similar Papers 제목 키워드 기반

iFLYTEK-Embodied-Omni Technical Report

2026-06-24 · Yuan Zhang, Jingfei Ni, Guanchen Lu, Shiqi Zhang 외 arxiv

General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. Existing approaches typically specialize …

Video Generation

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

2026-07-15 · Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo 외 hf

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoni…

Scene UnderstandingVisual ReasoningDecision Making

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

2025-10-20 · Yulin Luo, Chun-Kai Fan, Menghang Dong, Jiayu Shi 외 arxiv

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reason…

DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control

2024-07-20 · Xinyu Xu, Shengcheng Luo, Yanchao Yang, Yong-Lu Li 외

Building a general-purpose intelligent home-assistant agent skilled in diverse tasks by human commands is a long-term blueprint of embodied AI research, which poses requirements on task planning, environment modeling, an…

Instruction FollowingNavigateTask PlanningVision-Language Navigation

MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments

2024-02-01 · Yang Liu, Xinshuai Song, Kaixuan Jiang, Weixing Chen 외

With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically encode scene or historical memory in an u…

Embodied Question AnsweringLanguage ModelingLanguage ModellingLarge Language Model+2