paper-with-me

Papers

MonStereo: When Monocular and Stereo Meet at the Tail of 3D Human Localization

2020-08-25 · Lorenzo Bertoni, Sven Kreiss, Taylor Mordan, Alexandre Alahi

Monocular and stereo visions are cost-effective solutions for 3D human localization in the context of self-driving cars or social robots. However, they are usually developed independently and have their respective strengths and limitations. We propose a novel unified learning framework that leverages the strengths of both monocular and stereo cues for 3D human localization. Our method jointly (i) associates humans in left-right images, (ii) deals with occluded and distant cases in stereo settings by relying on the robustness of monocular cues, and (iii) tackles the intrinsic ambiguity of monocular perspective projection by exploiting prior knowledge of the human height distribution. We specifically evaluate outliers as well as challenging instances, such as occluded and far-away pedestrians, by analyzing the entire error distribution and by estimating calibrated confidence intervals. Finally, we critically review the official KITTI 3D metrics and propose a practical 3D localization metric tailored for humans.

📄 PDF Abstract BibTeX arXiv:2008.10913

Code (2)

vita-epfl/monstereo 공식 구현 pytorch
vita-epfl/monoloco pytorch

Tasks

Self-Driving Cars

Similar Papers 제목 키워드 기반

Robust Scale Estimation in Real-Time Monocular SFM for Autonomous Driving

2014-06-01 · CVPR 2014 6 · Shiyu Song, Manmohan Chandraker

Scale drift is a crucial challenge for monocular autonomous driving to emulate the performance of stereo. This paper presents a real-time monocular SFM system that corrects for scale drift using a novel cue combination f…

Autonomous DrivingObjectobject-detectionObject Detection+1

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

2026-07-06 · Jingyi Lu, Kai Han arxiv

Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bot…

Self-Supervised LearningVideo Generation

FusionDepth: Complement Self-Supervised Monocular Depth Estimation with Cost Volume

2023-05-10 · Zhuofei Huang, Jianlin Liu, Shang Xu, Ying Chen 외

Multi-view stereo depth estimation based on cost volume usually works better than self-supervised monocular depth estimation except for moving objects and low-textured surfaces. So in this paper, we propose a multi-frame…

Depth EstimationMonocular Depth EstimationStereo Depth Estimation

Diving into the Fusion of Monocular Priors for Generalized Stereo Matching

2025-05-20 · Chengtang Yao, Lidong Yu, Zhidan Liu, Jiaxi Zeng 외

The matching formulation makes it naturally hard for the stereo matching to handle ill-posed regions like occlusions and non-Lambertian surfaces. Fusing monocular priors has been proven helpful for ill-posed matching, bu…

Stereo Matching

SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input

2024-11-18 · CVPR 2025 1 · Zhen Lv, Yangqi Long, Congzhentao Huang, Cao Li 외

Stereo video synthesis from a monocular input is a demanding task in the fields of spatial computing and virtual reality. The main challenges of this task lie on the insufficiency of high-quality paired stereo videos for…

Novel View SynthesisVideo Generation