paper-with-me

홈 › Papers

Ctrl-VI: Controllable Video Synthesis via Variational Inference

2025-10-09 · Haoyi Duan, Yunzhi Zhang, Yilun Du, Jiajun Wu arxiv

Many video workflows benefit from a mixture of user controls with varying granularity, from exact 4D object trajectories and camera paths to coarse text prompts, while existing video generative models are typically trained for fixed input formats. We develop Ctrl-VI, a video synthesis method that addresses this need and generates samples with high controllability for specified elements while maintaining diversity for under-specified ones. We cast the task as variational inference to approximate a composed distribution, leveraging multiple video generation backbones to account for all task constraints collectively. To address the optimization challenge, we break down the problem into step-wise KL divergence minimization over an annealed sequence of distributions, and further propose a context-conditioned factorization technique that reduces modes in the solution space to circumvent local optima. Experiments suggest that our method produces samples with improved controllability, diversity, and 3D consistency compared to prior works.

📄 PDF Abstract BibTeX arXiv:2510.07670

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes

2025-05-30 · Anthony Gosselin, Ge Ya Luo, Luis Lara, Florian Golemo 외

Video diffusion techniques have advanced significantly in recent years; however, they struggle to generate realistic imagery of car crashes due to the scarcity of accident events in most driving datasets. Improving traff…

counterfactualVideo Generation

Ctrl-V: Higher Fidelity Video Generation with Bounding-Box Controlled Object Motion

2024-06-09 · Ge Ya Luo, Zhi Hao Luo, Anthony Gosselin, Alexia Jolicoeur-Martineau 외

Controllable video generation has attracted significant attention, largely due to advances in video diffusion models. In domains such as autonomous driving, it is essential to develop highly accurate predictions for obje…

Autonomous DrivingObjectVideo GenerationVideo Prediction

Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

2025-05-29 · Hengyuan Cao, Yutong Feng, Biao Gong, Yijing Tian 외

Video generative models can be regarded as world simulators due to their ability to capture dynamic, continuous changes inherent in real-world environments. These models integrate high-dimensional information across visu…

Dimensionality ReductionImage Generation

CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning

2024-10-15 · Qingqing Cao, Mahyar Najibi, Sachin Mehta

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising resul…

Image-text RetrievalText Retrievalzero-shot-classificationZero-Shot Learning

CtrlNeRF: The Generative Neural Radiation Fields for the Controllable Synthesis of High-fidelity 3D-Aware Images

2024-12-01 · Jian Liu, Zhen Yu

The neural radiance field (NERF) advocates learning the continuous representation of 3D geometry through a multilayer perceptron (MLP). By integrating this into a generative model, the generative neural radiance field (G…

3D geometryImage GenerationNeRF