paper-with-me

Papers

CRAFT: Video Diffusion for Bimanual Robot Data Generation

2026-04-04 · Jason Chen, I-Chun Arthur Liu, Gaurav Sukhatme, Daniel Seita arxiv

Bimanual robot learning from demonstrations is fundamentally limited by the cost and narrow visual diversity of real-world data, which constrains policy robustness across viewpoints, object configurations, and embodiments. We present Canny-guided Robot Data Generation using Video Diffusion Transformers (CRAFT), a video diffusion-based framework for scalable bimanual demonstration generation that synthesizes temporally coherent manipulation videos while producing action labels. By conditioning video diffusion on edge-based structural cues extracted from simulator-generated trajectories, CRAFT produces physically plausible trajectory variations and supports a unified augmentation pipeline spanning object pose changes, camera viewpoints, lighting and background variations, cross-embodiment transfer, and multi-view synthesis. We leverage a pre-trained video diffusion model to convert simulated videos, along with action labels from the simulation trajectories, into action-consistent demonstrations. Starting from only a few real-world demonstrations, CRAFT generates a large, visually diverse set of photorealistic training data, bypassing the need to replay demonstrations on the real robot (Sim2Real). Across simulated and real-world bimanual tasks, CRAFT improves success rates over existing augmentation strategies and straightforward data scaling, demonstrating that diffusion-based video generation can substantially expand demonstration diversity and improve generalization for dual-arm manipulation tasks. Our project website is available at: https://craftaug.github.io/

📄 PDF Abstract BibTeX arXiv:2604.03552

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Generalist Bimanual Manipulation via Foundation Video Diffusion Models

2025-07-17 · Yao Feng, Hengkai Tan, Xinyi Mao, Guodong Liu 외

Bimanual robotic manipulation, which involves the coordinated control of two robotic arms, is foundational for solving challenging tasks. Despite recent progress in general-purpose manipulation, data scarcity and embodim…

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

2024-10-10 · Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan 외

Bimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and th…

Zero-shot Generalization

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

2025-05-30 · Chenyou Fan, Fangzheng Yan, Chenjia Bai, Jiepeng Wang 외

Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existing approaches rely on Vision-Language-Act…

Action GenerationOptical Flow EstimationVideo PredictionVision-Language-Action

Diffusion-Based Imaginative Coordination for Bimanual Manipulation

2025-07-15 · Huilin Xu, Jian Ding, Jiakun Xu, Ruixiang Wang 외 arxiv

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordi…

Representation LearningVideo Prediction

SafeBimanual: Diffusion-based Trajectory Optimization for Safe Bimanual Manipulation

2025-08-25 · Haoyuan Deng, Wenkai Guo, Qianzhun Wang, Zhenyu Wu 외 arxiv

Bimanual manipulation has been widely applied in household services and manufacturing, which enables the complex task completion with coordination requirements. Recent diffusion-based policy learning approaches have achi…