paper-with-me

홈 › Papers

Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape Images

2025-04-04 · In-Hwan Jin, Haesoo Choo, Seong-Hun Jeong, Heemoon Park, Junghwan Kim, Oh-joon Kwon, Kyeongbo Kong

To achieve realistic immersion in landscape images, fluids such as water and clouds need to move within the image while revealing new scenes from various camera perspectives. Recently, a field called dynamic scene video has emerged, which combines single image animation with 3D photography. These methods use pseudo 3D space, implicitly represented with Layered Depth Images (LDIs). LDIs separate a single image into depth-based layers, which enables elements like water and clouds to move within the image while revealing new scenes from different camera perspectives. However, as landscapes typically consist of continuous elements, including fluids, the representation of a 3D space separates a landscape image into discrete layers, and it can lead to diminished depth perception and potential distortions depending on camera movement. Furthermore, due to its implicit modeling of 3D space, the output may be limited to videos in the 2D domain, potentially reducing their versatility. In this paper, we propose representing a complete 3D space for dynamic scene video by modeling explicit representations, specifically 4D Gaussians, from a single image. The framework is focused on optimizing 3D Gaussians by generating multi-view images from a single image and creating 3D motion to optimize 4D Gaussians. The most important part of proposed framework is consistent 3D motion estimation, which estimates common motion among multi-view images to bring the motion in 3D space closer to actual motions. As far as we know, this is the first attempt that considers animation while representing a complete 3D space from a single landscape image. Our model demonstrates the ability to provide realistic immersion in various landscape images through diverse experiments and metrics. Extensive experimental results are https://cvsp-lab.github.io/ICLR2025_3D-MOM/.

📄 PDF Abstract BibTeX arXiv:2504.05458

Code (1)

cvsp-lab/iclr2025_3d-mom 공식 구현 pytorch

Tasks

Image AnimationMotion Estimation

Similar Papers 제목 키워드 기반

3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint Videos

2024-03-03 · CVPR 2024 1 · Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang 외

Constructing photo-realistic Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos remains a challenging endeavor. Despite the remarkable advancements achieved by current neural rendering techniques, thes…

3DGSNeural Rendering

GUSH3R: Everyone Everywhere All at Once as Gaussians

2026-07-06 · Keito Abe, Kaede Shiohara, Takashi Otonari, Toshihiko Yamasaki arxiv

Reconstructing dynamic human-scene environments from monocular videos is a challenging problem that requires jointly modeling scene geometry, camera motion, and non-rigid human dynamics while enabling photorealistic rend…

Novel View Synthesis3D ReconstructionPoint Clouds

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

2025-01-01 · CVPR 2025 1 · Yuheng Jiang, Zhehao Shen, Chengcheng Guo, Yu Hong 외

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform gener…

AttributePosition

Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis

2023-08-18 · Jonathon Luiten, Georgios Kopanas, Bastian Leibe, Deva Ramanan

We present a method that simultaneously addresses the tasks of dynamic scene novel-view synthesis and six degree-of-freedom (6-DOF) tracking of all dense scene elements. We follow an analysis-by-synthesis framework, insp…

Dynamic ReconstructionNovel View SynthesisVideo Editing

Event-guided 3D Gaussian Splatting for Dynamic Human and Scene Reconstruction

2025-09-23 · Xiaoting Yin, Hao Shi, Kailun Yang, Jiajun Zhai 외 arxiv

Reconstructing dynamic humans together with static scenes from monocular videos remains difficult, especially under fast motion, where RGB frames suffer from motion blur. Event cameras exhibit distinct advantages, e.g., …