paper-with-me

Papers

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

2024-02-25 · Yao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao, Mingyu Ding, Ping Luo

Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these conceptual understandings into detailed robotic actions while achieving generalization across various scenarios. In this paper, we propose a tree-structured multimodal code generation framework for generalized robotic behavior synthesis, termed RoboCodeX. RoboCodeX decomposes high-level human instructions into multiple object-centric manipulation units consisting of physical preferences such as affordance and safety constraints, and applies code generation to introduce generalization ability across various robotics platforms. To further enhance the capability to map conceptual and perceptual understanding into control commands, a specialized multimodal reasoning dataset is collected for pre-training and an iterative self-updating methodology is introduced for supervised fine-tuning. Extensive experiments demonstrate that RoboCodeX achieves state-of-the-art performance in both simulators and real robots on four different kinds of manipulation tasks and one navigation task.

📄 PDF Abstract BibTeX arXiv:2402.16117

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationMultimodal ReasoningVisual Question Answering

Similar Papers 제목 키워드 기반

Behavior Generation with Latent Actions

2024-03-05 · Seungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim 외

Generative modeling of complex behaviors from labeled datasets has been a longstanding problem in decision making. Unlike language or image generation, decision making requires modeling actions - continuous-valued vector…

Autonomous DrivingDecision MakingImage GenerationImitation Learning+1

Multimodal Behavior Tree Generation: A Small Vision-Language Model for Robot Task Planning

2026-03-06 · Cristiano Battistini, Riccardo Andrea Izzo, Gianluca Bardaro, Matteo Matteucci arxiv

Large and small language models have been widely used for robotic task planning. At the same time, vision-language models (VLMs) have successfully tackled problems such as image captioning, scene understanding, and visua…

parameter-efficient fine-tuningVisual Question AnsweringRobot Task PlanningScene Understanding

PlayFusion: Skill Acquisition via Diffusion from Language-Annotated Play

2023-12-07 · Lili Chen, Shikhar Bahl, Deepak Pathak

Learning from unstructured and uncurated data has become the dominant paradigm for generative approaches in language and vision. Such unstructured and unguided behavior data, commonly known as play, is also easier to col…

Denoising

Ada3Drift: Adaptive Training-Time Drifting for One-Step 3D Visuomotor Robotic Manipulation

2026-03-12 · Chongyang Xu, Yixian Zou, Ziliang Feng, Fanman Meng 외 arxiv

Diffusion-based visuomotor policies effectively capture multimodal action distributions through iterative denoising, but their high inference latency limits real-time robotic control. Recent flow matching and consistency…

MARS Policy: Multimodality Only When It Matters

2026-05-28 · Jindou Jia, Tuo An, Yuxuan Hu, Gen Li 외 arxiv

Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavioral patterns, has driven the rapid emerge…

multimodal generation