GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
Despite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstruction and other metrically grounded downstream tasks. We propose GeometryCrafter, a novel framework that recovers high-fidelity point map sequences with temporal coherence from open-world videos, enabling accurate 3D/4D reconstruction, camera parameter estimation, and other depth-based applications. At the core of our approach lies a point map Variational Autoencoder (VAE) that learns a latent space agnostic to video latent distributions for effective point map encoding and decoding. Leveraging the VAE, we train a video diffusion model to model the distribution of point map sequences conditioned on the input videos. Extensive evaluations on diverse datasets demonstrate that GeometryCrafter achieves state-of-the-art 3D accuracy, temporal consistency, and generalization capability.
Code (0)
등록된 구현이 없습니다.
Tasks
4D reconstructionDepth Estimationparameter estimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
OMNI-PoseX: A Fast Vision Model for 6D Object Pose Estimation in Embodied Tasks
Accurate 6D object pose estimation is a fundamental capability for embodied agents, yet remains highly challenging in open-world environments. Many existing methods often rely on closed-set assumptions or geometry-agnost…
Zero-shot Generalization6D Pose EstimationRobotic GraspingOdometer-Agnostic Drift Correction Using OpenStreetMap Lane Geometry
Despite significant progress in odometry estimation, long-term drift remains a fundamental limitation of incremental pose integration, especially in large-scale or loop-free environments. Existing map-assisted methods ca…
Visual OdometryMatrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable…
Video GenerationRGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
3D Gaussian Splatting (3DGS) has become a popular solution in SLAM, as it can produce high-fidelity novel views. However, previous GS-based methods primarily target indoor scenes and rely on RGB-D sensors or pre-trained …
3DGSCamera Pose EstimationDepth EstimationNovel View Synthesis+2Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
Modern visual world modeling systems increasingly rely on high-capacity architectures and large-scale data to produce plausible motion, yet they often fail to preserve underlying 3D geometry or physically consistent came…
Depth Estimation