paper-with-me

홈 › Papers

DiVE: DiT-based Video Generation with Enhanced Control

2024-09-03 · Junpeng Jiang, Gangyi Hong, Lijun Zhou, Enhui Ma, Hengtong Hu, Xia Zhou, Jie Xiang, Fan Liu, Kaicheng Yu, Haiyang Sun, Kun Zhan, Peng Jia, Miao Zhang

Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent video generation works are proposed to tackcle the mentioned problem, i.e. models built on top of Diffusion Transformers (DiT), works are still missing which are targeted on exploring the potential for multi-view videos generation scenarios. Noticeably, we propose the first DiT-based framework specifically designed for generating temporally and multi-view consistent videos which precisely match the given bird's-eye view layouts control. Specifically, the proposed framework leverages a parameter-free spatial view-inflated attention mechanism to guarantee the cross-view consistency, where joint cross-attention modules and ControlNet-Transformer are integrated to further improve the precision of control. To demonstrate our advantages, we extensively investigate the qualitative comparisons on nuScenes dataset, particularly in some most challenging corner cases. In summary, the effectiveness of our proposed method in producing long, controllable, and highly consistent videos under difficult conditions is proven to be effective.

📄 PDF Abstract BibTeX arXiv:2409.01595

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Real-Time Motion-Controllable Autoregressive Video Diffusion

2025-10-09 · Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou 외 arxiv

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion model…

Text-to-Video GenerationReinforcement Learning

SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model

2024-09-05 · Weipeng Tan, Chuming Lin, Chengming Xu, Xiaozhong Ji 외

Talking Head Generation (THG), typically driven by audio, is an important and challenging task with broad application prospects in various fields such as digital humans, film production, and virtual reality. While diffus…

DiversityTalking Head Generation

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

2025-10-16 · Yuancheng Xu, Wenqi Xian, Li Ma, Julien Philip 외 arxiv

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with r…

Video Generation

DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis

2026-05-16 · Yi Zuo, Huimin Wu, Lingling Li, Fang Liu 외 arxiv

Trajectory-controlled video generation has become essential for controllable video generation. While current methods perform well under small-view camera motions, they degrade significantly with large-view motions. Exist…

Video Generation

CustomX: Unified Character, Action, and Scene Customization in Video World Models

2025-12-18 · Yitong Wang, Fangyun Wei, Hongyang Zhang, Bo Dai 외 arxiv

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without acti…

Video Generation