paper-with-me

홈 › Papers

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

2025-03-18 · Nvidia, :, Hassan Abu Alhaija, Jose Alvarez, Maciej Bala, Tiffany Cai, Tianshi Cao, Liz Cha, Joshua Chen, Mike Chen, Francesco Ferroni, Sanja Fidler, Dieter Fox, Yunhao Ge, Jinwei Gu, Ali Hassani, Michael Isaev, Pooya Jannaty, Shiyi Lan, Tobias Lasser, Huan Ling, Ming-Yu Liu, Xian Liu, Yifan Lu, Alice Luo, Qianli Ma, Hanzi Mao, Fabio Ramos, Xuanchi Ren, Tianchang Shen, Xinglong Sun, Shitao Tang, Ting-Chun Wang, Jay Wu, Jiashu Xu, Stella Xu, Kevin Xie, Yuchong Ye, Xiaodong Yang, Xiaohui Zeng, Yu Zeng

We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, the spatial conditional scheme is adaptive and customizable. It allows weighting different conditional inputs differently at different spatial locations. This enables highly controllable world generation and finds use in various world-to-world transfer use cases, including Sim2Real. We conduct extensive evaluations to analyze the proposed model and demonstrate its applications for Physical AI, including robotics Sim2Real and autonomous vehicle data enrichment. We further demonstrate an inference scaling strategy to achieve real-time world generation with an NVIDIA GB200 NVL72 rack. To help accelerate research development in the field, we open-source our models and code at https://github.com/nvidia-cosmos/cosmos-transfer1.

📄 PDF Abstract BibTeX arXiv:2503.14492

Code (2)

nvidia-cosmos/cosmos-transfer1 공식 구현 pytorch
nv-tlabs/cosmos-drive-dreams pytorch

Similar Papers 제목 키워드 기반

World Simulation with Video Foundation Models for Physical AI

2025-10-28 · NVIDIA, :, Arslan Ali, Junjie Bai 외 arxiv

We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World gene…

Synthetic Data GenerationReinforcement LearningVideo Generation

Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

2025-06-10 · Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao 외

Collecting and annotating real-world data for safety-critical physical AI systems, such as Autonomous Vehicle (AV), is time-consuming and costly. It is especially challenging to capture rare edge cases, which play a crit…

3D Lane Detection3D Object DetectionLane Detectionobject-detection+3

Cosmos 3: Omnimodal World Models for Physical AI

2026-06-01 · NVIDIA, :, Aditi, Niket Agarwal 외 arxiv

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting …

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

2026-01-22 · Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin 외 arxiv

Recent video generation models demonstrate remarkable ability to capture complex physical interactions and scene evolution over time. To leverage their spatiotemporal priors, robotics works have adapted video models for …

Video Generation

Cosmos World Foundation Model Platform for Physical AI

2025-01-07 · Nvidia, :, Niket Agarwal, Arslan Ali 외

Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present the Cosmos World Foundation Model Platform…

modelPosition