paper-with-me

홈 › Papers

Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss

2018-12-08 · Lipu Zhou, Jiamin Ye, Montiel Abello, Shengze Wang, Michael Kaess

We present a novel unsupervised learning framework for single view depth estimation using monocular videos. It is well known in 3D vision that enlarging the baseline can increase the depth estimation accuracy, and jointly optimizing a set of camera poses and landmarks is essential. In previous monocular unsupervised learning frameworks, only part of the photometric and geometric constraints within a sequence are used as supervisory signals. This may result in a short baseline and overfitting. Besides, previous works generally estimate a low resolution depth from a low resolution impute image. The low resolution depth is then interpolated to recover the original resolution. This strategy may generate large errors on object boundaries, as the depth of background and foreground are mixed to yield the high resolution depth. In this paper, we introduce a bundle adjustment framework and a super-resolution network to solve the above two problems. In bundle adjustment, depths and poses of an image sequence are jointly optimized, which increases the baseline by establishing the relationship between farther frames. The super resolution network learns to estimate a high resolution depth from a low resolution image. Additionally, we introduce the clip loss to deal with moving objects and occlusion. Experimental results on the KITTI dataset show that the proposed algorithm outperforms the state-of-the-art unsupervised methods using monocular sequences, and achieves comparable or even better result compared to unsupervised methods using stereo sequences.

📄 PDF Abstract BibTeX arXiv:1812.03368

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMonocular Depth EstimationSuper-Resolution

Similar Papers 제목 키워드 기반

Marginalized Bundle Adjustment: Multi-View Camera Pose from Monocular Depth Estimates

2026-02-21 · Shengjie Zhu, Ahmed Abdelkader, Mark J. Matthews, Xiaoming Liu 외 arxiv

Structure-from-Motion (SfM) is a fundamental 3D vision task for recovering camera parameters and scene geometry from multi-view images. While recent deep learning advances enable accurate Monocular Depth Estimation (MDE)…

Monocular Depth EstimationPoint Clouds

Dynamic Visual SLAM using a General 3D Prior

2025-12-07 · Xingguang Zhong, Liren Jin, Marija Popović, Jens Behley 외 arxiv

Reliable incremental estimation of camera poses and 3D reconstruction is key to enable various applications including robotics, interactive visualization, and augmented reality. However, this task is particularly challen…

Camera Pose Estimation3D Reconstruction

GlORIE-SLAM: Globally Optimized RGB-only Implicit Encoding Point Cloud SLAM

2024-03-28 · Ganlin Zhang, Erik Sandström, Youmin Zhang, Manthan Patel 외

Recent advancements in RGB-only dense Simultaneous Localization and Mapping (SLAM) have predominantly utilized grid-based neural implicit encodings and/or struggle to efficiently realize global map and pose consistency. …

Simultaneous Localization and Mapping

PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment

2026-06-30 · Aymeric Fleith, Julian Zirbel, Daniel Cremers, Niclas Zeller arxiv

We introduce PRISM-VO, a novel pure optimization-based sparse photometric visual odometry framework for focused plenoptic cameras. The core of PRISM-VO is a novel photometric plenoptic bundle adjustment which jointly opt…

Visual Odometry

Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual Odometry

2024-06-03 · Takayuki Kanai, Igor Vasiljevic, Vitor Guizilini, Kazuhiro Shintani

Monocular visual odometry is a key technology in a wide variety of autonomous systems. Relative to traditional feature-based methods, that suffer from failures due to poor lighting, insufficient texture, large motions, e…

Depth EstimationMonocular Depth EstimationMonocular Visual OdometryVisual Odometry