paper-with-me

Papers

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

2024-12-04 · CVPR 2025 1 · Jiahao Lu, Tianyu Huang, Peng Li, Zhiyang Dou, Cheng Lin, Zhiming Cui, Zhen Dong, Sai-Kit Yeung, Wenping Wang, YuAn Liu

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the input video, which is training-expensive and can only produce scale-invariant depth values without camera poses. In this paper, we propose a novel video-depth estimation method called Align3R to estimate temporal consistent depth maps for a dynamic video. Our key idea is to utilize the recent DUSt3R model to align estimated monocular depth maps of different timesteps. First, we fine-tune the DUSt3R model with additional estimated monocular depth as inputs for the dynamic scenes. Then, we apply optimization to reconstruct both depth maps and camera poses. Extensive experiments demonstrate that Align3R estimates consistent video depth and camera poses for a monocular video with superior performance than baseline methods.

📄 PDF Abstract BibTeX arXiv:2412.03079

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Self-supervised Event-based Monocular Depth Estimation using Cross-modal Consistency

2024-01-14 · Junyu Zhu, Lina Liu, Bofeng Jiang, Feng Wen 외

An event camera is a novel vision sensor that can capture per-pixel brightness changes and output a stream of asynchronous ``events''. It has advantages over conventional cameras in those scenes with high-speed motions a…

Depth EstimationDepth PredictionMonocular Depth Estimation

MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network

2025-07-15 · Jianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu 외

Learning-based Multi-View Stereo (MVS) methods aim to predict depth maps for a sequence of calibrated images to recover dense point clouds. However, existing MVS methods often struggle with challenging regions, such as t…

Depth EstimationDepth PredictionMonocular Depth Estimation

Learning Monocular Depth in Dynamic Environment via Context-aware Temporal Attention

2023-05-12 · Zizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu 외

The monocular depth estimation task has recently revealed encouraging prospects, especially for the autonomous driving task. To tackle the ill-posed problem of 3D geometric reasoning from 2D monocular images, multi-frame…

Autonomous DrivingDepth EstimationMonocular Depth EstimationPose Estimation

Focusable Monocular Depth Estimation

2026-05-12 · Yuxin Du, Tao Lin, Zile Zhong, Runting Li 외 arxiv

Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-relevant target regions from the surroun…

Monocular Depth Estimation

3D Human Pose Estimation via Explicit Compositional Depth Maps

2020-02-08 · AAAI 2020 2 · Haiping Wu, Bin Xiao

n this work, we tackle the problem of estimating 3D human pose in camera space from a monocular image. First, we propose to use densely-generated limb depth maps to ease the learning of body joints depth, which are well …

3D Human Pose EstimationPose Estimation