paper-with-me

홈 › Papers

Can Video Diffusion Model Reconstruct 4D Geometry?

2025-03-27 · Jinjie Mai, Wenxuan Zhu, Haozhe Liu, Bing Li, Cheng Zheng, Jürgen Schmidhuber, Bernard Ghanem

Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods either require specialized 4D representation or sophisticated optimization. In this paper, we present Sora3R, a novel framework that taps into the rich spatiotemporal priors of large-scale video diffusion models to directly infer 4D pointmaps from casual videos. Sora3R follows a two-stage pipeline: (1) we adapt a pointmap VAE from a pretrained video VAE, ensuring compatibility between the geometry and video latent spaces; (2) we finetune a diffusion backbone in combined video and pointmap latent space to generate coherent 4D pointmaps for every frame. Sora3R operates in a fully feedforward manner, requiring no external modules (e.g., depth, optical flow, or segmentation) or iterative global alignment. Extensive experiments demonstrate that Sora3R reliably recovers both camera poses and detailed scene geometry, achieving performance on par with state-of-the-art methods for dynamic 4D reconstruction across diverse scenarios.

📄 PDF Abstract BibTeX arXiv:2503.21082

Code (0)

등록된 구현이 없습니다.

Tasks

4D reconstructionmodelOptical Flow Estimation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors

2025-04-01 · Tian-Xing Xu, Xiangjun Gao, WenBo Hu, Xiaoyu Li 외

Despite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstru…

4D reconstructionDepth Estimationparameter estimation

VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling

2026-06-12 · Xunzhi Xiang, Zixuan Duan, Yabo Chen, Zhengxuan Wei 외 arxiv

Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually enforce geometric consistency by using ex…

Video GenerationPoint Clouds

GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction

2026-02-27 · Chao Xu, Xiaochen Zhao, Xiang Deng, Jingxiang Sun 외 arxiv

Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video gen…

Video Generation

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models

2026-04-12 · Dehui Wang, Rong Wei, Yue Shi, Congsheng Xu 외 arxiv

The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches struggle to infer massive amounts of missing…

Video Super-ResolutionScene Generation

Spatio-Temporal Garment Reconstruction Using Diffusion Mapping via Pattern Coordinates

2026-02-27 · Yingxuan You, Ren Li, Corentin Dumery, Cong Cao 외 arxiv

Reconstructing 3D clothed humans from monocular images and videos is a fundamental problem with applications in virtual try-on, avatar creation, and mixed reality. Despite significant progress in human body recovery, acc…

Dynamic ReconstructionGarment ReconstructionVirtual Try-on