paper-with-me

홈 › Papers

Learning Depth from Monocular Videos Using Synthetic Data: A Temporally-Consistent Domain Adaptation Approach

2019-07-16 · Yipeng Mou, Mingming Gong, Huan Fu, Kayhan Batmanghelich, Kun Zhang, DaCheng Tao

Majority of state-of-the-art monocular depth estimation methods are supervised learning approaches. The success of such approaches heavily depends on the high-quality depth labels which are expensive to obtain. Some recent methods try to learn depth networks by leveraging unsupervised cues from monocular videos which are easier to acquire but less reliable. In this paper, we propose to resolve this dilemma by transferring knowledge from synthetic videos with easily obtainable ground-truth depth labels. Due to the stylish difference between synthetic and real images, we propose a temporally-consistent domain adaptation (TCDA) approach that simultaneously explores labels in the synthetic domain and temporal constraints in the videos to improve style transfer and depth prediction. Furthermore, we make use of the ground-truth optical flow and pose information in the synthetic data to learn moving mask and pose prediction networks. The learned moving masks can filter out moving regions that produces erroneous temporal constraints and the estimated poses provide better initializations for estimating temporal constraints. Experimental results demonstrate the effectiveness of our method and comparable performance against state-of-the-art.

📄 PDF Abstract BibTeX arXiv:1907.06882

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationDepth PredictionDomain AdaptationMonocular Depth EstimationOptical Flow EstimationPose PredictionStyle Transfer

Similar Papers 제목 키워드 기반

ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors

2025-09-16 · Romain Hardy, Tyler Berzin, Pranav Rajpurkar arxiv

Three-dimensional (3D) scene understanding in colonoscopy presents significant challenges that necessitate automated methods for accurate depth estimation. However, existing depth estimation models for endoscopy struggle…

Point Cloud GenerationScene Understanding3D ReconstructionDepth Estimation

Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting

2025-04-15 · Jiaxin Huang, Sheng Miao, Bangbang Yang, Yuewen Ma 외

Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D …

4D reconstructionVideo Inpainting

GenMM: Geometrically and Temporally Consistent Multimodal Data Generation for Video and LiDAR

2024-06-15 · Bharat Singh, Viveka Kulharia, Luyu Yang, Avinash Ravichandran 외

Multimodal synthetic data generation is crucial in domains such as autonomous driving, robotics, augmented/virtual reality, and retail. We propose a novel approach, GenMM, for jointly editing RGB videos and LiDAR scans b…

Autonomous DrivingDepth EstimationMonocular Depth EstimationSemantic Segmentation+2

Seurat: From Moving Points to Depth

2025-04-20 · CVPR 2025 1 · Seokju Cho, Jiahui Huang, Seungryong Kim, Joon-Young Lee

Accurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However, humans often perceive relative depth int…

Depth EstimationPoint Tracking

GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion

2025-05-29 · Gwanghyun Kim, Xueting Li, Ye Yuan, Koki Nagano 외

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies…

Depth EstimationImage to Video GenerationVideo Generation