Deconstructing Self-Supervised Monocular Reconstruction: The Design Decisions that Matter
This paper presents an open and comprehensive framework to systematically evaluate state-of-the-art contributions to self-supervised monocular depth estimation. This includes pretraining, backbone, architectural design choices and loss functions. Many papers in this field claim novelty in either architecture design or loss formulation. However, simply updating the backbone of historical systems results in relative improvements of 25%, allowing them to outperform the majority of existing systems. A systematic evaluation of papers in this field was not straightforward. The need to compare like-with-like in previous papers means that longstanding errors in the evaluation protocol are ubiquitous in the field. It is likely that many papers were not only optimized for particular datasets, but also for errors in the data and evaluation criteria. To aid future research in this area, we release a modular codebase (https://github.com/jspenmar/monodepth_benchmark), allowing for easy evaluation of alternate design decisions against corrected data and evaluation criteria. We re-implement, validate and re-evaluate 16 state-of-the-art contributions and introduce a new dataset (SYNS-Patches) containing dense outdoor depth maps in a variety of both natural and urban scenes. This allows for the computation of informative metrics in complex regions such as depth boundaries.
Code (2)
Tasks
Depth EstimationMonocular Depth EstimationMonocular ReconstructionSimilar Papers 제목 키워드 기반
MonoSelfRecon: Purely Self-Supervised Explicit Generalizable 3D Reconstruction of Indoor Scenes from Monocular RGB Views
Current monocular 3D scene reconstruction (3DR) works are either fully-supervised, or not generalizable, or implicit in 3D representation. We propose a novel framework - MonoSelfRecon that for the first time achieves exp…
3D Reconstruction3D Scene ReconstructionDepth EstimationNeRFMonocular Depth Estimation with Self-supervised Instance Adaptation
Recent advances in self-supervised learning havedemonstrated that it is possible to learn accurate monoculardepth reconstruction from raw video data, without using any 3Dground truth for supervision. However, in robotics…
Depth EstimationMonocular Depth EstimationMonocular ReconstructionSelf-Supervised LearningSelf-supervised Dense 3D Reconstruction from Monocular Endoscopic Video
We present a self-supervised learning-based pipeline for dense 3D reconstruction from full-length monocular endoscopic videos without a priori modeling of anatomy or shading. Our method only relies on unlabeled monocular…
3D ReconstructionAnatomySelf-Supervised LearningSynthesizing Light Field Video from Monocular Video
The hardware challenges associated with light-field(LF) imaging has made it difficult for consumers to access its benefits like applications in post-capture focus and aperture control. Learning-based techniques which sol…
Self-Supervised LearningVideo ReconstructionS2F2: Self-Supervised High Fidelity Face Reconstruction from Monocular Image
We present a novel face reconstruction method capable of reconstructing detailed face geometry, spatially varying face reflectance from a single monocular image. We build our work upon the recent advances of DNN-based au…
3D Face ReconstructionFace ReconstructionSelf-Supervised LearningVocal Bursts Intensity Prediction