paper-with-me

홈 › Papers

Object-Centric Instruction Augmentation for Robotic Manipulation

2024-01-05 · Junjie Wen, Yichen Zhu, Minjie Zhu, Jinming Li, Zhiyuan Xu, Zhengping Che, Chaomin Shen, Yaxin Peng, Dong Liu, Feifei Feng, Jian Tang

Humans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as \enquote{pick and place}, understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the literature that uses the large language model to enrich the text descriptions, the latter remains underexplored. In this work, we introduce the \textit{Object-Centric Instruction Augmentation (OCI)} framework to augment highly semantic and information-dense language instruction with position cues. We utilize a Multi-modal Large Language Model (MLLM) to weave knowledge of object locations into natural language instruction, thus aiding the policy network in mastering actions for versatile manipulation. Additionally, we present a feature reuse mechanism to integrate the vision-language features from off-the-shelf pre-trained MLLM into policy networks. Through a series of simulated and real-world robotic tasks, we demonstrate that robotic manipulator imitation policies trained with our enhanced instructions outperform those relying solely on traditional language instructions.

📄 PDF Abstract BibTeX arXiv:2401.02814

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelObject

Similar Papers 제목 키워드 기반

LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation

2026-03-26 · Motonari Kambara, Koki Seno, Tomoya Kaichi, Yanan Wang 외 arxiv

We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This …

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow

2025-07-08 · Yixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang 외

Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate this issue, they remain limited to scenar…

Deformable Object ManipulationImitation LearningObject

Object-Centric World Model for Language-Guided Manipulation

2025-03-08 · Youngjoon Jeong, Junha Chun, Soonwoo Cha, Taesup Kim

A world model is essential for an agent to predict the future and plan in domains such as autonomous driving and robotics. To achieve this, recent advancements have focused on video generation, which has gained significa…

Autonomous DrivingmodelObjectObject Recognition+1

STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation

2026-01-28 · Alexandre Chapin, Emmanuel Dellandréa, Liming Chen arxiv

Visual foundation models provide strong perceptual features for robotics, but their dense representations lack explicit object-level structure, limiting robustness and controllability in manipulation tasks. We propose ST…

FOCUS: Object-Centric World Models for Robotics Manipulation

2023-07-05 · Stefano Ferraro, Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt

Understanding the world in terms of objects and the possible interplays with them is an important cognition ability, especially in robotics manipulation, where many tasks require robot-object interactions. However, learn…

Object