paper-with-me

홈 › Papers

GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth

2026-05-11 · Yuecheng Liu, Junda Cheng, Longliang Liu, Wenjing Liao, Hanrui Cheng, Yuzhou Wang, Xin Yang arxiv

Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail regions and temporal inconsistencies. We argue that current approaches, which primarily rely on temporal smoothing via Transformers, struggle to maintain strict 3D geometric consistency-particularly under rotations or drastic view changes. To address this, we propose GemDepth, a framework built on the insight that an explicit awareness of camera motion and global 3D structure is a prerequisite for 3D consistency. Distinctively, GemDepth introduces a Geometry-Embedding Module (GEM) that predicts inter-frame camera poses to generate implicit geometric embeddings. This injection of motion priors equips the network with intrinsic 3D perception and alignment capabilities. Guided by these geometric cues, our Alternating Spatio-Temporal Transformer (ASTT) captures latent point-level correspondences to simultaneously enhance spatial precision for sharp details and enforce rigorous temporal consistency. Furthermore, GemDepth employs a data-efficient training strategy, effectively bridging the gap between high efficiency and robust geometric consistency. As shown in Fig.2, comprehensive evaluations demonstrate that GemDepth achieves state-of-the-art performance across multiple datasets, particularly in complex dynamic scenarios. The code is publicly available at: https://github.com/Yuecheng919/GemDepth.

📄 PDF Abstract BibTeX arXiv:2605.10525

Code (0)

등록된 구현이 없습니다.

Tasks

Depth Estimation

Similar Papers 제목 키워드 기반

LEGO: Learning Edge with Geometry all at Once by Watching Videos

2018-03-15 · CVPR 2018 6 · Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu 외

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a "3D as-smooth-as-possible (3D-ASAP)" prior…

3D geometryAll

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

2025-07-10 · Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo 외 arxiv

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in …

Video Generation

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

2025-07-03 · Fangfu Liu, Hao Li, Jiawei Chi, Hanyang Wang 외 arxiv

Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent developments have achieved this by performing per-scene optimization with embedded language info…

Scene Understanding

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

2026-03-16 · Minjun Kang, Inkyu Shin, Taeyeop Lee, Myungchul Kim 외 arxiv

Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, …

Novel View Synthesis

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

2026-05-14 · Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi 외 arxiv

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do n…

Video Generation