paper-with-me

Papers

Prediction with Action: Visual Policy Learning via Joint Denoising Process

2024-11-27 · Yanjiang Guo, Yucheng Hu, Jianke Zhang, Yen-Jen Wang, Xiaoyu Chen, Chaochao Lu, Jianyu Chen

Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical world. On the other line, diffusion models have also shown promise in robotic control tasks by denoising actions, known as diffusion policy. Although the diffusion generative model and diffusion policy exhibit distinct capabilities--image prediction and robotic action, respectively--they technically follow a similar denoising process. In robotic tasks, the ability to predict future images and generate actions is highly correlated since they share the same underlying dynamics of the physical world. Building on this insight, we introduce PAD, a novel visual policy learning framework that unifies image Prediction and robot Action within a joint Denoising process. Specifically, PAD utilizes Diffusion Transformers (DiT) to seamlessly integrate images and robot states, enabling the simultaneous prediction of future images and robot actions. Additionally, PAD supports co-training on both robotic demonstrations and large-scale video datasets and can be easily extended to other robotic modalities, such as depth images. PAD outperforms previous methods, achieving a significant 26.3% relative improvement on the full Metaworld benchmark, by utilizing a single text-conditioned visual policy within a data-efficient imitation learning setting. Furthermore, PAD demonstrates superior generalization to unseen tasks in real-world robot manipulation settings with 28.0% success rate increase compared to the strongest baseline. Project page at https://sites.google.com/view/pad-paper

📄 PDF Abstract BibTeX arXiv:2411.18179

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage GenerationImitation LearningRobot Manipulation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models

2026-03-19 · Zirui Ge, Pengxiang Ding, Baohua Yin, Qishen Wang 외 arxiv

Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge to downstream robot control. Yet current…

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation

2026-06-02 · Dingrui Wang, YuAn Wang, Jinkun Liu, Yue Zhang 외 arxiv

Recent world action models leverage video foundation models by aligning broad visual-dynamics priors with executable robot actions. We revisit this alignment from a distributional perspective. Existing formulations typic…

Point Tracking Improves World Action Models

2026-05-22 · Jiarui Guan, Wenshuai Zhao, Yue Pei, Ziliang Chen 외 arxiv

Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and texture, making learned representations …

Point Tracking

Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

2024-12-27 · Yongbiao Gao, Xiangcheng Sun, Guohua Lv, Deng Yu 외

Audio-visual video parsing (AVVP) aims to recognize audio and visual event labels with precise temporal boundaries, which is quite challenging since audio or visual modality might include only one event label with only t…

Denoising

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

2026-09-11 · Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim 외 hf

Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and…

Trajectory Modeling