paper-with-me

Papers

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

2025-10-16 · Yuancheng Xu, Wenqi Xian, Li Ma, Julien Philip, Ahmet Levent Taşel, Yiwei Zhao, Ryan Burgert, Mingming He, Oliver Hermann, Oliver Pilarski, Rahul Garg, Paul Debevec, Ning Yu arxiv

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric capture performances re-rendered with diverse camera trajectories via 4D Gaussian Splatting (4DGS), lighting variability obtained with a video relighting model. We fine-tune state-of-the-art open-source video diffusion models on this data to provide strong multi-view identity preservation, precise camera control, and lighting adaptability. Our framework also supports core capabilities for virtual production, including multi-subject generation using two approaches: joint training and noise blending, the latter enabling efficient composition of independently customized models at inference time; it also achieves scene and real-life video customization as well as control over motion and spatial layout during customization. Extensive experiments show improved video quality, higher personalization accuracy, and enhanced camera control and lighting adaptability, advancing the integration of video generation into virtual production. Our project page is available at: https://eyeline-labs.github.io/Virtually-Being.

📄 PDF Abstract BibTeX arXiv:2510.14179

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

2026-04-10 · Haoyu Zhao, Zihao Zhang, Jiaxi Gu, Haoran Chen 외 arxiv

Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either provide imprecise camera control from text prompts or rely on labor…

Spatial ReasoningVideo Generation

One-shot lip-based biometric authentication: extending behavioral features with authentication phrase information

2023-08-14 · Brando Koch, Ratko Grbić

Lip-based biometric authentication (LBBA) is an authentication method based on a person's lip movements during speech in the form of video data captured by a camera sensor. LBBA can utilize both physical and behavioral c…

One-Shot LearningTriplet

MotionMaster: Training-free Camera Motion Transfer For Video Generation

2024-04-24 · Teng Hu, Jiangning Zhang, Ran Yi, Yating Wang 외

The emergence of diffusion models has greatly propelled the progress in image and video generation. Recently, some efforts have been made in controllable video generation, including text-to-video generation and video mot…

DisentanglementMotion DisentanglementText-to-Video GenerationVideo Generation

Automatic Camera Control and Directing with an Ultra-High-Definition Collaborative Recording System

2022-08-10 · Bram Vanherle, Tim Vervoort, Nick Michiels, Philippe Bekaert

Capturing an event from multiple camera angles can give a viewer the most complete and interesting picture of that event. To be suitable for broadcasting, a human director needs to decide what to show at each point in ti…

object-detectionObject Detection

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

2024-07-17 · Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace 외

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is crit…

Video Generation