paper-with-me

홈 › Papers

Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics

2026-03-11 · Tianshuo Xu, Zhifei Chen, Leyi Wu, Hao Lu, Ying-cong Chen arxiv

The ultimate goal of video generation is to satisfy a fundamental trilemma: achieving high visual quality, maintaining rigorous physical consistency, and enabling precise controllability. While recent models can maintain this balance in simple, isolated scenarios, we observe that this equilibrium is fragile and often breaks down as scene complexity increases (e.g., involving collisions or dense traffic). To address this, we introduce \textbf{Motion Forcing}, a framework designed to stabilize this trilemma even in complex generative tasks. Our key insight is to explicitly decouple physical reasoning from visual synthesis via a hierarchical \textbf{``Point-Shape-Appearance''} paradigm. This approach decomposes generation into verifiable stages: modeling complex dynamics as sparse geometric anchors (\textbf{Point}), expanding them into dynamic depth maps that explicitly resolve 3D geometry (\textbf{Shape}), and finally rendering high-fidelity textures (\textbf{Appearance}). Furthermore, to foster robust physical understanding, we employ a \textbf{Masked Point Recovery} strategy. By randomly masking input anchors during training and enforcing the reconstruction of complete dynamic depth, the model is compelled to move beyond passive pattern matching and learn latent physical laws (e.g., inertia) to infer missing trajectories. Extensive experiments on autonomous driving benchmarks show that Motion Forcing significantly outperforms state-of-the-art baselines, maintaining trilemma stability across complex scenes. Evaluations on physics and robotics further confirm our framework's generality.

📄 PDF Abstract BibTeX arXiv:2603.10408

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVideo Generation

Similar Papers 제목 키워드 기반

Decoupled Self-Forcing Distillation for Streaming Talking Head Generation

2026-09-09 · Yanru An, Ruiyan Wang, Wenwu Wei, Rui Bu 외 arxiv

Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directions: end-to-end methods condition a video diffusion model on audio dir…

Talking Head Generation

X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio

2025-08-04 · Chenxu Zhang, Zenan Li, Hongyi Xu, You Xie 외 arxiv

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that e…

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

2026-05-09 · Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu 외 arxiv

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often …

Video Generation

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

2025-12-20 · Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 외 arxiv

With the rise of online dance-video platforms and rapid advances in AI-generated content (AIGC), music-driven dance generation has emerged as a compelling research direction. Despite substantial progress in related domai…

Video Generation

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

2024-12-08 · CVPR 2025 1 · Shuwei Shi, Biao Gong, Xi Chen, Dandan Zheng 외

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate divers…

Contrastive LearningImage to Video GenerationMotion EstimationOptical Flow Estimation+2