paper-with-me

홈 › Papers

Prism: Semi-Supervised Multi-View Stereo with Monocular Structure Priors

2024-12-08 · Alex Rich, Noah Stier, Pradeep Sen, Tobias Höllerer

The promise of unsupervised multi-view-stereo (MVS) is to leverage large unlabeled datasets, yet current methods underperform when training on difficult data, such as handheld smartphone videos of indoor scenes. Meanwhile, high-quality synthetic datasets are available but MVS networks trained on these datasets fail to generalize to real-world examples. To bridge this gap, we propose a semi-supervised learning framework that allows us to train on real and rendered images jointly, capturing structural priors from synthetic data while ensuring parity with the real-world domain. Central to our framework is a novel set of losses that leverages powerful existing monocular relative-depth estimators trained on the synthetic dataset, transferring the rich structure of this relative depth to the MVS predictions on unlabeled data. Inspired by perceptual image metrics, we compare the MVS and monocular predictions via a deep feature loss and a multi-scale statistical loss. Our full framework, which we call Prism, achieves large quantitative and qualitative improvements over current unsupervised and synthetic-supervised MVS networks. This is a best-case-scenario result, opening the door to using both unlabeled smartphone videos and photorealistic synthetic datasets for training MVS networks.

📄 PDF Abstract BibTeX arXiv:2412.05771

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Semi-Supervised Stereo-Based 3D Object Detection via Cross-View Consensus

2023-01-01 · CVPR 2023 1 · Wenhao Wu, Hau San Wong, Si Wu

Stereo-based 3D object detection, which aims at detecting 3D objects with stereo cameras, shows great potential in low-cost deployment compared to LiDAR-based methods and excellent performance compared to monocular-b…

3D Object DetectionDepth EstimationObjectobject-detection+2

LentiAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic Communication

2026-06-09 · Chufeng Fang, Dongdong Teng, Lilin Liu arxiv

Real-time stereoscopic video communication has long been a goal of immersive telepresence, yet practical systems still require specialized capture rigs or reduce remote users to a single portrait view. We present LentiAv…

Semi-supervised Deep Multi-view Stereo

2022-07-24 · Hongbin Xu, Weitao Chen, Yang Liu, Zhipeng Zhou 외

Significant progress has been witnessed in learning-based Multi-view Stereo (MVS) under supervised and unsupervised settings. To combine their respective merits in accuracy and completeness, meantime reducing the demand …

Semi-Supervised Monocular Depth Estimation with Left-Right Consistency Using Deep Neural Network

2019-05-18 · Ali Jahani Amiri, Shing Yan Loo, Hong Zhang

There has been tremendous research progress in estimating the depth of a scene from a monocular camera image. Existing methods for single-image depth prediction are exclusively based on deep neural networks, and their tr…

Depth EstimationDepth PredictionMonocular Depth EstimationPrediction

MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments from a Single Moving Camera

2020-11-24 · CVPR 2021 1 · Felix Wimbauer, Nan Yang, Lukas von Stumberg, Niclas Zeller 외

In this paper, we propose MonoRec, a semi-supervised monocular dense reconstruction architecture that predicts depth maps from a single moving camera in dynamic environments. MonoRec is based on a multi-view stereo setti…