paper-with-me

홈 › Papers

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

2025-06-06 · Hongyan Zhi, Peihao Chen, Siyuan Zhou, Yubo Dong, Quanxi Wu, Lei Han, Mingkui Tan

Manipulation has long been a challenging task for robots, while humans can effortlessly perform complex interactions with objects, such as hanging a cup on the mug rack. A key reason is the lack of a large and uniform dataset for teaching robots manipulation skills. Current robot datasets often record robot action in different action spaces within a simple scene. This hinders the robot to learn a unified and robust action representation for different robots within diverse scenes. Observing how humans understand a manipulation task, we find that understanding how the objects should move in the 3D space is a critical clue for guiding actions. This clue is embodiment-agnostic and suitable for both humans and different robots. Motivated by this, we aim to learn a 3D flow world model from both human and robot manipulation data. This model predicts the future movement of the interacting objects in 3D space, guiding action planning for manipulation. Specifically, we synthesize a large-scale 3D optical flow dataset, named ManiFlow-110k, through a moving object auto-detect pipeline. A video diffusion-based world model then learns manipulation physics from these data, generating 3D optical flow trajectories conditioned on language instructions. With the generated 3D object optical flow, we propose a flow-guided rendering mechanism, which renders the predicted final state and leverages GPT-4o to assess whether the predicted flow aligns with the task description. This equips the robot with a closed-loop planning ability. Finally, we consider the predicted 3D optical flow as constraints for an optimization policy to determine a chunk of robot actions for manipulation. Extensive experiments demonstrate strong generalization across diverse robotic manipulation tasks and reliable cross-embodiment adaptation without hardware-specific training.

📄 PDF Abstract BibTeX arXiv:2506.06199

Code (1)

hoyyyaard/3dflowaction 공식 구현 pytorch

Tasks

Optical Flow EstimationRobot Manipulation

Similar Papers 제목 키워드 기반

ContactFlow: A video action conditioning that transfers across embodiments

2026-07-29 · Sami Azirar, Enrico Pallotta, Jan Nogga, Jürgen Gall 외 arxiv

World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the ph…

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow

2025-07-08 · Yixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang 외

Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate this issue, they remain limited to scenar…

Deformable Object ManipulationImitation LearningObject

Trajectory Conditioned Cross-embodiment Skill Transfer

2025-10-09 · YuHang Tang, Yixuan Lou, Pengfei Han, Haoming Song 외 arxiv

Learning manipulation skills from human demonstration videos presents a promising yet challenging problem, primarily due to the significant embodiment gap between human body and robot manipulators. Existing methods rely …

Robot ManipulationTransfer Learning

Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields for Flow-Based Object Manipulation

2026-06-22 · Koki Seno, Daichi Yashima, Yusuke Takagi, Kento Tokura 외 arxiv

Cross-embodiment data have become central to training robotic foundation models. To leverage such heterogeneous data, we focus on flow-based object manipulation, where robot flows (robot velocity fields) serve as embodim…

GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

2025-08-14 · Kelin Yu, Sheng Zhang, Harshit Soora, Furong Huang 외 arxiv

Recent advances have shown that video generation models can enhance robot learning by deriving effective robot actions through inverse dynamics. However, these methods heavily depend on the quality of generated data and …

Reinforcement LearningVideo Generation