paper-with-me

홈 › Papers

MAD: Motion Appearance Decoupling for efficient Driving World Models

2026-01-14 · Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord, Alexandre Alahi arxiv

Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are essential. Adapting these generalist video models to driving domains has shown promise but typically requires massive domain-specific data and costly fine-tuning. We propose an efficient adaptation framework that converts generalist video diffusion models into controllable driving world models with minimal supervision. The key idea is to decouple motion learning from appearance synthesis. First, the model is adapted to predict structured motion in a simplified form: videos of skeletonized agents and scene elements, focusing learning on physical and social plausibility. Then, the same backbone is reused to synthesize realistic RGB videos conditioned on these motion sequences, effectively "dressing" the motion with texture and lighting. This two-stage process mirrors a reasoning-rendering paradigm: first infer dynamics, then render appearance. Our experiments show this decoupled approach is exceptionally efficient: adapting SVD, we match prior SOTA models with less than 6% of their compute. Scaling to LTX, our MAD-LTX model outperforms all open-source competitors, and supports a comprehensive suite of text, ego, and object controls. Project page: https://vita-epfl.github.io/MAD-World-Model/

📄 PDF Abstract BibTeX arXiv:2601.09452

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

CDNet: A cascaded decoupling architecture for video prediction

2021-09-29 · Chuanqi Zang, Mingtao Pei

Video prediction is an essential task in the computer vision community, helping to solve many downstream vision tasks by predicting and modeling future motion dynamics and appearance. In the deterministic video predictio…

PredictionVideo Prediction

Image Animation with Keypoint Mask

2021-12-20 · Or Toledano, Yanir Marmor, Dov Gertz

Motion transfer is the task of synthesizing future video frames of a single source image according to the motion from a given driving video. This task is challenging due to the complexity of motion representation and the…

Image Animation

GA-Drive: Geometry-Appearance Decoupled Modeling for Free-viewpoint Driving Scene Generation

2026-02-24 · Hao Zhang, Lue Fan, Qitai Wang, Wenbo Li 외 arxiv

A free-viewpoint, editable, and high-fidelity driving simulator is crucial for training and evaluating end-to-end autonomous driving systems. In this paper, we present GA-Drive, a novel simulation framework capable of ge…

Autonomous DrivingScene Generation

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

2025-03-24 · CVPR 2025 1 · Zunnan Xu, Zhentao Yu, Zixiang Zhou, Jun Zhou 외

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance refer…

DenoisingPortrait Animation

Extending Multi-Object Tracking systems to better exploit appearance and 3D information

2019-12-25 · Kanchana Ranasinghe, Sahan Liyanaarachchi, Harsha Ranasinghe, Mayuka Jayawardhana

Tracking multiple objects in real time is essential for a variety of real-world applications, with self-driving industry being at the foremost. This work involves exploiting temporally varying appearance and motion infor…

Multi-Object TrackingObjectObject TrackingReal-Time Multi-Object Tracking