paper-with-me

Papers

DisCo: World Models with Discrete Camera Motion Control

2026-06-06 · Hongrui Huang, Junke Wang, Quanhao Li, Yu-Gang Jiang, Zuxuan Wu arxiv

Controllable video world models target interactive world exploration, where models must faithfully execute explicit action commands while preserving visual quality and temporal coherence. However, most existing approaches rely on continuous camera trajectories as action conditions, which often lead to unreliable action following, especially under complex motion sequences. In this work, we identify action representation entanglement as a key bottleneck in controllable video generation, and show that continuous camera representations lead to high feature similarity across distinct motion patterns, degrading action controllability. Based on this insight, we propose DisCo, a controllable video world model that conditions generation on a compact set of discrete action primitives to improve action separability. We further introduce DisCoBench, a comprehensive benchmark for evaluating the ability of models in short-term, long-horizon, and highly dynamic exploration scenarios. Extensive experiments demonstrate that DisCo achieves significantly more reliable action following while preserving visual quality.

📄 PDF Abstract BibTeX arXiv:2606.07967

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

EchoWM: Open and Enterable Omnimodal World Models

2026-08-24 · Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin 외 arxiv

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around…

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

2026-06-26 · Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang 외 arxiv

Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accu…

Video Generation

Wonder: Video World Model Done Better

2026-07-28 · Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli 외 arxiv

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactivel…

Video Generation

HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

2026-07-02 · Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu, Xintao Wang 외 arxiv

We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand annotations from multi-view or marker-based…

Video Generation

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

2026-04-04 · Guiyu Zhang, Yabo Chen, Xunzhi Xiang, Junchao Huang 외 arxiv

Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle only one motion type or rely on ambiguous 2D cues that entangle camera-in…

Video Generation