paper-with-me

Papers

Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video

2026-06-27 · Qing Zhao, Weijian Deng, Pengxu Wei, Liang Lin arxiv

Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over time remains challenging. Dynamic Gaussian Splatting achieves strong rendering performance through photometric optimization, yet does not explicitly enforce multi-view geometric consistency. In contrast, 3D foundation models recover coherent scene geometry and camera motion, but their point-based outputs are not designed for photorealistic rendering. We propose Ground4D, a geometry-grounded framework built on two stages. First, we perform geometry initialization via 3D foundation models, leveraging VGGT in a training-free manner to reconstruct multi-view-consistent 3D geometry and camera poses from monocular video. The recovered geometry provides a structured and reliable initialization for dynamic Gaussian representations. Second, we conduct geometry-consistency-aware refinement via dynamic Gaussian Splatting, optimizing the representation through differentiable rendering while maintaining multi-view geometric consistency across both observed and synthesized viewpoints. Furthermore, Ground4D inherently models the continuous 4D dynamics of the scene, naturally supporting rendering at arbitrary timestamps. By integrating foundation-level geometric priors into dynamic Gaussian optimization, Ground4D achieves stronger reconstruction fidelity and rendering performance, underscoring the role of geometry-grounded constraints in robust 4D scene modeling.

📄 PDF Abstract BibTeX arXiv:2606.28828

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gravity-Aware Monocular 3D Human-Object Reconstruction

2021-08-19 · ICCV 2021 10 · Rishabh Dabral, Soshi Shimada, Arjun Jain, Christian Theobalt 외

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free…

Human-Object Interaction DetectionObjectObject Reconstruction

Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting

2024-06-04 · Inkyu Shin, Qihang Yu, Xiaohui Shen, In So Kweon 외

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address this, we introduce Video-3DGS, a 3D Gauss…

3DGSNeRFVideo EditingVideo Reconstruction

Temporal Consistency Loss for High Resolution Textured and Clothed 3DHuman Reconstruction from Monocular Video

2021-04-19 · Akin Caliskan, Armin Mustafa, Adrian Hilton

We present a novel method to learn temporally consistent 3D reconstruction of clothed people from a monocular video. Recent methods for 3D human reconstruction from monocular video using volumetric, implicit or parametri…

3D geometry3D Human Reconstruction3D Human Shape Estimation3D Reconstruction+1

CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction

2026-05-20 · Hao Xu, Yilin Liu, Yinqiao Wang, Chi-Wing Fu 외 arxiv

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. Such a capability wo…

Endo3R: Unified Online Reconstruction from Dynamic Monocular Endoscopic Video

2025-04-04 · Jiaxin Guo, Wenzhen Dong, Tianyu Huang, Hao Ding 외

Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction rem…

Camera Pose EstimationDepth EstimationDepth PredictionDynamic Reconstruction+1