paper-with-me

Papers

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

2025-04-10 · Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic prior captured by such video models, Geo4D can be trained using only synthetic data while generalizing well to real data in a zero-shot manner. Geo4D predicts several complementary geometric modalities, namely point, depth, and ray maps. It uses a new multi-modal alignment algorithm to align and fuse these modalities, as well as multiple sliding windows, at inference time, thus obtaining robust and accurate 4D reconstruction of long videos. Extensive experiments across multiple benchmarks show that Geo4D significantly surpasses state-of-the-art video depth estimation methods, including recent methods such as MonST3R, which are also designed to handle dynamic scenes.

📄 PDF Abstract BibTeX arXiv:2504.07961

Code (1)

jzr99/Geo4D 공식 구현 pytorch

Tasks

3D Reconstruction4D reconstructionDepth Estimation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction

2026-01-07 · Jiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma 외 arxiv

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric la…

Scene Generation3D GenerationPoint Clouds

ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

2026-03-04 · Zihao Huang, Tianqi Liu, Zhaoxi Chen, Shaocong Xu 외 arxiv

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage video diffusion models to synthesize hum…

Inverse Rendering

World Reconstruction From Inconsistent Views

2026-03-17 · Lukas Höllein, Matthias Nießner arxiv

Video diffusion models generate high-quality and diverse worlds; however, individual frames often lack 3D consistency across the output sequence, which makes the reconstruction of 3D worlds difficult. To this end, we pro…

3D Reconstruction

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

2025-10-15 · Hyojun Go, Dominik Narnhofer, Goutam Bhat, Prune Truong 외 arxiv

The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could obtain a formidable 3D scene generator …

3D Reconstruction3D Generation

MonoNeuralFusion: Online Monocular Neural 3D Reconstruction with Geometric Priors

2022-09-30 · Zi-Xin Zou, Shi-Sheng Huang, Yan-Pei Cao, Tai-Jiang Mu 외

High-fidelity 3D scene reconstruction from monocular videos continues to be challenging, especially for complete and fine-grained geometry reconstruction. The previous 3D reconstruction approaches with neural implicit re…

3D Reconstruction3D Scene Reconstruction