paper-with-me

Papers

Editable Free-viewpoint Video Using a Layered Neural Representation

2021-04-30 · Jiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao, Yanshun Zhang, Minye Wu, Yingliang Zhang, Lan Xu, Jingyi Yu

Generating free-viewpoint videos is critical for immersive VR/AR experience but recent neural advances still lack the editing ability to manipulate the visual perception for large dynamic scenes. To fill this gap, in this paper we propose the first approach for editable photo-realistic free-viewpoint video generation for large-scale dynamic scenes using only sparse 16 cameras. The core of our approach is a new layered neural representation, where each dynamic entity including the environment itself is formulated into a space-time coherent neural layered radiance representation called ST-NeRF. Such layered representation supports fully perception and realistic manipulation of the dynamic scene whilst still supporting a free viewing experience in a wide range. In our ST-NeRF, the dynamic entity/layer is represented as continuous functions, which achieves the disentanglement of location, deformation as well as the appearance of the dynamic entity in a continuous and self-supervised manner. We propose a scene parsing 4D label map tracking to disentangle the spatial information explicitly, and a continuous deform module to disentangle the temporal motion implicitly. An object-aware volume rendering scheme is further introduced for the re-assembling of all the neural layers. We adopt a novel layered loss and motion-aware ray sampling strategy to enable efficient training for a large dynamic scene with multiple performers, Our framework further enables a variety of editing functions, i.e., manipulating the scale and location, duplicating or retiming individual neural layers to create numerous visual effects while preserving high realism. Extensive experiments demonstrate the effectiveness of our approach to achieve high-quality, photo-realistic, and editable free-viewpoint video generation for dynamic scenes.

📄 PDF Abstract BibTeX arXiv:2104.14786

Code (1)

darlinghang/st-nerf pytorch

Tasks

DisentanglementNeRFScene ParsingVideo Generation

Similar Papers 제목 키워드 기반

GA-Drive: Geometry-Appearance Decoupled Modeling for Free-viewpoint Driving Scene Generation

2026-02-24 · Hao Zhang, Lue Fan, Qitai Wang, Wenbo Li 외 arxiv

A free-viewpoint, editable, and high-fidelity driving simulator is crucial for training and evaluating end-to-end autonomous driving systems. In this paper, we present GA-Drive, a novel simulation framework capable of ge…

Autonomous DrivingScene Generation

MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines

2026-03-03 · Ryan Po, David Junhao Zhang, Amir Hertz, Gordon Wetzstein 외 arxiv

Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user control over the environment for reproducib…

4D-MoDe: Towards Editable and Scalable Volumetric Streaming via Motion-Decoupled 4D Gaussian Compression

2025-09-22 · Houqiang Zhong, Zihan Zheng, Qiang Hu, Yuan Tian 외 arxiv

Volumetric video has emerged as a key medium for immersive telepresence and augmented/virtual reality, enabling six-degrees-of-freedom (6DoF) navigation and realistic spatial interactions. However, delivering high-qualit…

The Visual Centrifuge: Model-Free Layered Video Representations

2018-12-04 · CVPR 2019 6 · Jean-Baptiste Alayrac, João Carreira, Andrew Zisserman

True video understanding requires making sense of non-lambertian scenes where the color of light arriving at the camera sensor encodes information about not just the last object it collided with, but about multiple mediu…

Color ConstancymodelVideo Understanding

PVP: Personalized Video Prior for Editable Dynamic Portraits using StyleGAN

2023-06-29 · Kai-En Lin, Alex Trevithick, Keli Cheng, Michel Sarkis 외

Portrait synthesis creates realistic digital avatars which enable users to interact with others in a compelling way. Recent advances in StyleGAN and its extensions have shown promising results in synthesizing photorealis…

Face Generation