paper-with-me

홈 › Papers

Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

2026-06-30 · Zixing Wang, Kausik Sivakumar, Jinghuan Shang, Yafei Hu, Zhaoming Xie, Ran Gong, Xiaohan Zhang, Karl Schmeckpeper arxiv

Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet world-action models have not previously been shown to transfer from simulation to real robotic manipulation. We study whether a world-action model can be trained from synthetic priors and deployed zero-shot in the real world. To this end, we build upon Cosmos Policy, a video diffusion model adapted for visuomotor control. We construct simulation environments with extensive domain randomization and generate demonstrations using the AnyTask motion planning pipeline. We evaluate our approach across object lifting, drawer opening, and pick-and-place tasks using ${\sim}800$ synthetic demonstrations per task and no real demonstrations. When deployed zero-shot on a Franka Robot, our policy attains a 35\% average success rate. To our knowledge, this represents the first successful sim-to-real transfer of a world-action model for robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2606.31101

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Planning

Similar Papers 제목 키워드 기반

Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation

2026-03-17 · Xinhao Cai, Gensheng Pei, Zeren Sun, Yazhou Yao 외 arxiv

In this paper, we propose \textbf{Iris}, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional feed-forward methods rely on massive traini…

Monocular Depth Estimation

Learning Transferable Dynamics Priors from Action to World Modeling

2026-06-28 · Ze Huang, Jiahui Zhang, Hairuo Liu, Chenxi Zhang 외 arxiv

We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the resulting world model…

Robot ManipulationVideo Generation

EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision

2026-05-13 · Jiahao Chen, Zihui Zhang, Yafei Yang, Jinxi Li 외 arxiv

We introduce EvObj for unsupervised 3D instance segmentation that bridges the geometric domain gap between synthetic pretraining data and real-world point clouds. Current methods suffer from structural discrepancies when…

3D Instance SegmentationObject SegmentationPoint Clouds

Learning 3D Robotics Perception using Inductive Priors

2024-05-30 · Muhammad Zubair Irshad

Recent advances in deep learning have led to a data-centric intelligence i.e. artificially intelligent models unlocking the potential to ingest a large amount of data and be really good at performing digital tasks such a…

3D ReconstructionImage GenerationInductive BiasScene Understanding+3

Prism: Semi-Supervised Multi-View Stereo with Monocular Structure Priors

2024-12-08 · Alex Rich, Noah Stier, Pradeep Sen, Tobias Höllerer

The promise of unsupervised multi-view-stereo (MVS) is to leverage large unlabeled datasets, yet current methods underperform when training on difficult data, such as handheld smartphone videos of indoor scenes. Meanwhil…