paper-with-me

홈 › Papers

FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation

2025-05-15 · Jun Guo, Xiaojian Ma, Yikai Wang, Min Yang, Huaping Liu, Qing Li

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canonical approaches that handle dynamics prediction mostly implicitly and reconcile it with visual rendering in a single model, we introduce FlowDreamer, which adopts 3D scene flow as explicit motion representations. FlowDreamer first predicts 3D scene flow from past frame and action conditions with a U-Net, and then a diffusion model will predict the future frame utilizing the scene flow. FlowDreamer is trained end-to-end despite its modularized nature. We conduct experiments on 4 different benchmarks, covering both video prediction and visual planning tasks. The results demonstrate that FlowDreamer achieves better performance compared to other baseline RGB-D world models by 7% on semantic similarity, 11% on pixel quality, and 6% on success rate in various robot manipulation domains.

📄 PDF Abstract BibTeX arXiv:2505.10075

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationSemantic SimilaritySemantic Textual SimilarityVideo Prediction

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields for Flow-Based Object Manipulation

2026-06-22 · Koki Seno, Daichi Yashima, Yusuke Takagi, Kento Tokura 외 arxiv

Cross-embodiment data have become central to training robotic foundation models. To leverage such heterogeneous data, we focus on flow-based object manipulation, where robot flows (robot velocity fields) serve as embodim…

FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow

2024-08-09 · Hangyu Li, Xiangxiang Chu, Dingyuan Shi, Wang Lin

Recent advances in text-to-3D generation have made significant progress. In particular, with the pretrained diffusion models, existing methods predominantly use Score Distillation Sampling (SDS) to train 3D models such a…

3D GenerationNeRFText to 3D

Object-centric 3D Motion Field for Robot Learning from Human Videos

2025-06-04 · Zhao-Heng Yin, Sherry Yang, Pieter Abbeel

Learning robot control policies from human videos is a promising direction for scaling up robot learning. However, how to extract action knowledge (or action representations) from videos for policy learning remains a key…

DenoisingMotion Estimation

Robotic VLA Benefits from Joint Learning with Motion Image Diffusion

2025-12-19 · Yu Fang, Kanchana Ranasinghe, Le Xue, Honglu Zhou 외 arxiv

Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they typically mimic expert trajectories wit…

Future Optical Flow Prediction Improves Robot Control & Video Generation

2026-01-15 · Kanchana Ranasinghe, Honglu Zhou, Yu Fang, Luyu Yang 외 arxiv

Future motion representations, such as optical flow, offer immense value for control and generative tasks. However, forecasting generalizable spatially dense motion representations remains a key challenge, and learning s…

Multimodal ReasoningVideo Generation