paper-with-me

홈 › Papers

Ctrl-World: A Controllable Generative World Model for Robot Manipulation

2025-10-11 · Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea Finn arxiv

Generalist robot policies can now perform a wide range of manipulation skills, but evaluating and improving their ability with unfamiliar objects and instructions remains a significant challenge. Rigorous evaluation requires a large number of real-world rollouts, while systematic improvement demands additional corrective data with expert labels. Both of these processes are slow, costly, and difficult to scale. World models offer a promising, scalable alternative by enabling policies to rollout within imagination space. However, a key challenge is building a controllable world model that can handle multi-step interactions with generalist robot policies. This requires a world model compatible with modern generalist policies by supporting multi-view prediction, fine-grained action control, and consistent long-horizon interactions, which is not achieved by previous works. In this paper, we make a step forward by introducing a controllable multi-view world model that can be used to evaluate and improve the instruction-following ability of generalist robot policies. Our model maintains long-horizon consistency with a pose-conditioned memory retrieval mechanism and achieves precise action control through frame-level action conditioning. Trained on the DROID dataset (95k trajectories, 564 scenes), our model generates spatially and temporally consistent trajectories under novel scenarios and new camera placements for over 20 seconds. We show that our method can accurately rank policy performance without real-world robot rollouts. Moreover, by synthesizing successful trajectories in imagination and using them for supervised fine-tuning, our approach can improve policy success by 44.7\%.

📄 PDF Abstract BibTeX arXiv:2510.10125

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

2025-05-29 · Hengyuan Cao, Yutong Feng, Biao Gong, Yijing Tian 외

Video generative models can be regarded as world simulators due to their ability to capture dynamic, continuous changes inherent in real-world environments. These models integrate high-dimensional information across visu…

Dimensionality ReductionImage Generation

CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning

2024-03-29 · Luke Rowe, Roger Girgis, Anthony Gosselin, Bruno Carrez 외

Evaluating autonomous vehicle stacks (AVs) in simulation typically involves replaying driving logs from real-world recorded traffic. However, agents replayed from offline data are not reactive and hard to intuitively con…

counterfactualOffline RLreinforcement-learningReinforcement Learning+1

Controlling Neural Networks with Rule Representations

2021-06-14 · NeurIPS 2021 12 · Sungyong Seo, Sercan O. Arik, Jinsung Yoon, Xiang Zhang 외

We propose a novel training method that integrates rules into deep learning, in a way the strengths of the rules are controllable at inference. Deep Neural Networks with Controllable Rule Representations (DeepCTRL) incor…

Decision Making

Semantically Controllable Augmentations for Generalizable Robot Learning

2024-09-02 · Zoey Chen, Zhao Mandi, Homanga Bharadhwaj, Mohit Sharma 외

Generalization to unseen real-world scenarios for robot manipulation requires exposure to diverse datasets during training. However, collecting large real-world datasets is intractable due to high operational costs. For …

Data AugmentationRobot Manipulation

SweCTRL-Mini: a data-transparent Transformer-based large language model for controllable text generation in Swedish

2023-04-27 · Dmytro Kalpakchi, Johan Boye

We present SweCTRL-Mini, a large Swedish language model that can be used for inference and fine-tuning on a single consumer-grade GPU. The model is based on the CTRL architecture by Keskar, McCann, Varshney, Xiong, and S…

GPULanguage ModelingLanguage ModellingLarge Language Model+1