paper-with-me

Papers

X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering

2025-10-09 · Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang, Hao Zhu, Jing Liao arxiv

We present X2Video, the first diffusion model for rendering photorealistic videos guided by intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, while supporting intuitive multi-modal controls with reference images and text prompts for both global and local regions. The intrinsic guidance allows accurate manipulation of color, material, geometry, and lighting, while reference images and text prompts provide intuitive adjustments in the absence of intrinsic information. To enable these functionalities, we extend the intrinsic-guided image generation model XRGB to video generation by employing a novel and efficient Hybrid Self-Attention, which ensures temporal consistency across video frames and also enhances fidelity to reference images. We further develop a Masked Cross-Attention to disentangle global and local text prompts, applying them effectively onto respective local and global regions. For generating long videos, our novel Recursive Sampling method incorporates progressive frame sampling, combining keyframe prediction and frame interpolation to maintain long-range temporal consistency while preventing error accumulation. To support the training of X2Video, we assembled a video dataset named InteriorVideo, featuring 1,154 rooms from 295 interior scenes, complete with reliable ground-truth intrinsic channel sequences and smooth camera trajectories. Both qualitative and quantitative evaluations demonstrate that X2Video can produce long, temporally consistent, and photorealistic videos guided by intrinsic conditions. Additionally, X2Video effectively accommodates multi-modal controls with reference images, global and local text prompts, and simultaneously supports editing on color, material, geometry, and lighting through parametric tuning. Project page: https://luckyhzt.github.io/x2video

📄 PDF Abstract BibTeX arXiv:2510.08530

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationImage Generation

Similar Papers 제목 키워드 기반

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

2026-05-28 · Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou 외 arxiv

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models re…

Video Generation

M3-CVC: Controllable Video Compression with Multimodal Generative Models

2024-11-24 · Rui Wan, Qi Zheng, Yibo Fan

Traditional and neural video codecs commonly encounter limitations in controllability and generality under ultra-low-bitrate coding scenarios. To overcome these challenges, we propose M3-CVC, a controllable video compres…

Video Compression

HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models

2024-06-20 · Xinrui Zhou, Yuhao Huang, Wufeng Xue, Haoran Dou 외

Echocardiography (ECHO) video is widely used for cardiac examination. In clinical, this procedure heavily relies on operator experience, which needs years of training and maybe the assistance of deep learning-based syste…

Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

2024-01-03 · David Junhao Zhang, Dongxu Li, Hung Le, Mike Zheng Shou 외

Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the generated videos. This work presents Moon…

Image AnimationVideo EditingVideo Generation

AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction

2024-06-10 · Zhen Xing, Qi Dai, Zejia Weng, Zuxuan Wu 외

Text-guided video prediction (TVP) involves predicting the motion of future frames from the initial frame according to an instruction, which has wide applications in virtual reality, robotics, and content creation. Previ…

Language ModellingLarge Language ModelVideo Prediction