paper-with-me

홈 › Papers

LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

2024-12-12 · Yabo Chen, Chen Yang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie, Wei Shen, Wenrui Dai, Hongkai Xiong, Qi Tian

Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer promising 3D priors learned from large-scale video data. However, leveraging these priors effectively faces three key challenges: (1) degradation in quality across large camera motions, (2) difficulties in achieving precise camera control, and (3) geometric distortions inherent to the diffusion process that damage 3D consistency. We address these challenges by proposing LiftImage3D, a framework that effectively releases LVDMs' generative priors while ensuring 3D consistency. Specifically, we design an articulated trajectory strategy to generate video frames, which decomposes video sequences with large camera motions into ones with controllable small motions. Then we use robust neural matching models, i.e. MASt3R, to calibrate the camera poses of generated frames and produce corresponding point clouds. Finally, we propose a distortion-aware 3D Gaussian splatting representation, which can learn independent distortions between frames and output undistorted canonical Gaussians. Extensive experiments demonstrate that LiftImage3D achieves state-of-the-art performance on two challenging datasets, i.e. LLFF, DL3DV, and Tanks and Temples, and generalizes well to diverse in-the-wild images, from cartoon illustrations to complex real-world scenes.

📄 PDF Abstract BibTeX arXiv:2412.09597

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionImage to 3DVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Dream, Lift, Animate: From Single Images to Animatable Gaussian Avatars

2025-07-21 · Marcel C. Bühler, Ye Yuan, Xueting Li, Yangyi Huang 외 arxiv

We introduce Dream, Lift, Animate (DLA), a novel framework that reconstructs animatable 3D human avatars from a single image. This is achieved by leveraging multi-view generation, 3D Gaussian lifting, and pose-aware UV-s…

Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing

2025-08-05 · Hongyu Shen, Junfeng Ni, Yixin Chen, Weishuo Li 외 arxiv

We address the challenge of lifting 2D visual segmentation to 3D in Gaussian Splatting. Existing methods often suffer from inconsistent 2D masks across viewpoints and produce noisy segmentation boundaries as they neglect…

4K4DGen: Panoramic 4D Generation at 4K Resolution

2024-06-19 · Renjie Li, Panwang Pan, Bangbang Yang, Dejia Xu 외

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques ei…

4k

Generalizable and Animatable Gaussian Head Avatar

2024-10-10 · Xuangeng Chu, Tatsuya Harada

In this paper, we propose Generalizable and Animatable Gaussian head Avatar (GAGAvatar) for one-shot animatable head avatar reconstruction. Existing methods rely on neural radiance fields, leading to heavy rendering cons…

Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction

2026-05-29 · Mungyeom Kim, Minkyeong Jeon, Honggyu An, Jaewoo Jung 외 arxiv

Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame, suffering from duplicated Gaussians and …

Scene UnderstandingPoint Tracking