paper-with-me

홈 › Papers

Pixel Motion Diffusion is What We Need for Robot Control

2025-09-26 · E-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe, Xiang Li, Michael S. Ryoo arxiv

We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion intent and low-level robot action via structured pixel motion representation. In DAWN, both the high-level and low-level controllers are modeled as diffusion processes, yielding a fully trainable, end-to-end system with interpretable intermediate motion abstractions. DAWN achieves state-of-the-art results on the challenging CALVIN benchmark, demonstrating strong multi-task performance, and further validates its effectiveness on MetaWorld. Despite the substantial domain gap between simulation and reality and limited real-world data, we demonstrate reliable real-world transfer with only minimal finetuning, illustrating the practical viability of diffusion-based motion abstractions for robotic control. Our results show the effectiveness of combining diffusion modeling with motion-centric representations as a strong baseline for scalable and robust robot learning. Project page: https://eronguyen.github.io/DAWN/

📄 PDF Abstract BibTeX arXiv:2509.22652

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pixel Motion as Universal Representation for Robot Control

2025-05-12 · Kanchana Ranasinghe, Xiang Li, Cristina Mata, Jongwoo Park 외

We present LangToMo, a vision-language-action framework structured as a dual-system architecture that uses pixel motion forecasts as intermediate representations. Our high-level System 2, an image diffusion model, genera…

Vision-Language-Action

Point Tracking Improves World Action Models

2026-05-22 · Jiarui Guan, Wenshuai Zhao, Yue Pei, Ziliang Chen 외 arxiv

Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and texture, making learned representations …

Point Tracking

Robotic VLA Benefits from Joint Learning with Motion Image Diffusion

2025-12-19 · Yu Fang, Kanchana Ranasinghe, Le Xue, Honglu Zhou 외 arxiv

Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they typically mimic expert trajectories wit…

DeepBlindness: Fast Blindness Map Estimation and Blindness Type Classification for Outdoor Scene from Single Color Image

2019-11-02 · Jiaxiong Qiu, Xinyuan Yu, Guoqiang Yang, Shuaicheng Liu

Outdoor vision robotic systems and autonomous cars suffer from many image-quality issues, particularly haze, defocus blur, and motion blur, which we will define generically as "blindness issues". These blindness issues m…

General Classification

DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy

2025-06-25 · Sungjae Park, Homanga Bharadhwaj, Shubham Tulsiani

We propose DemoDiffusion, a simple and scalable method for enabling robots to perform manipulation tasks in natural environments by imitating a single human demonstration. Our approach is based on two key insights. First…