paper-with-me

Papers

BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

2026-02-03 · Yixiang Chen, Peiyan Li, Jiabing Yang, Keji He, Xiangnan Wu, Yuan Xu, Kai Wang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang arxiv

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual and motion priors. However, they still face key challenges: a misalignment between coordinate-space actions and pixel-space videos, sensitivity to camera viewpoint, and non-unified architectures across embodiments. To this end, we present BridgeV2W, which converts coordinate-space actions into pixel-aligned embodiment masks rendered from the URDF and camera parameters. These masks are then injected into a pretrained video generation model via a ControlNet-style pathway, which aligns the action control signals with predicted videos, adds view-specific conditioning to accommodate camera viewpoints, and yields a unified world model architecture across embodiments. To mitigate overfitting to static backgrounds, BridgeV2W further introduces a flow-based motion loss that focuses on learning dynamic and task-relevant regions. Experiments on single-arm (DROID) and dual-arm (AgiBot-G1) datasets, covering diverse and challenging conditions with unseen viewpoints and scenes, show that BridgeV2W improves video generation quality compared to prior state-of-the-art methods. We further demonstrate the potential of BridgeV2W on downstream real-world tasks, including policy evaluation and goal-conditioned planning. More results can be found on our project website at https://BridgeV2W.github.io .

📄 PDF Abstract BibTeX arXiv:2602.03793

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

2026-08-05 · Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma 외 hf

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, ex…

Robot ManipulationPoint Clouds

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

2024-12-16 · Qi Sun, Pengfei Hong, Tej Deep Pala, Vernon Toh 외

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate st…

HallucinationRobot ManipulationScene UnderstandingSpatial Reasoning+1

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

2026-03-30 · Muyang He, Hanzhong Guo, Junxiong Lin, Yizhou Yu arxiv

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between…

Autonomous DrivingVideo Generation

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

2026-07-13 · Xinghang Li, Jun Guo, Qiwei Li, Long Qian 외 hf

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coh…

Text-to-Image GenerationScene GenerationVideo GenerationImage Editing

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

2025-11-27 · Quanjian Song, Yiren Song, Kelly Peng, Yuan Gao 외 arxiv

Recent advances in video world models enable interactive environments with free navigation, making translation between first-person (egocentric) and third-person (exocentric) perspectives increasingly important. However,…

Video Generation