Depth Is All You Need for Monocular 3D Detection
A key contributor to recent progress in 3D detection from single images is monocular depth estimation. Existing methods focus on how to leverage depth explicitly, by generating pseudo-pointclouds or providing attention cues for image features. More recent works leverage depth prediction as a pretraining task and fine-tune the depth representation while training it for 3D detection. However, the adaptation is insufficient and is limited in scale by manual labels. In this work, we propose to further align depth representation with the target domain in unsupervised fashions. Our methods leverage commonly available LiDAR or RGB videos during training time to fine-tune the depth representation, which leads to improved 3D detectors. Especially when using RGB videos, we show that our two-stage training by first generating pseudo-depth labels is critical because of the inconsistency in loss distribution between the two tasks. With either type of reference data, our multi-task learning approach improves over the state of the art on both KITTI and NuScenes, while matching the test-time complexity of its single task sub-network.
Code (0)
등록된 구현이 없습니다.
Tasks
AllDepth EstimationDepth PredictionMonocular Depth EstimationMulti-Task LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Is Pseudo-Lidar needed for Monocular 3D Object detection?
Recent progress in 3D object detection from single images leverages monocular depth estimation as a way to produce 3D pointclouds, turning cameras into pseudo-lidar sensors. These two-stage detectors improve with the acc…
3D Object DetectionDepth EstimationMonocular 3D Object DetectionMonocular Depth Estimation+3Probabilistic and Geometric Depth: Detecting Objects in Perspective
3D object detection is an important capability needed in various practical applications such as driver assistance systems. Monocular 3D detection, as a representative general setting among image-based approaches, provide…
3D Object DetectionAttributeDepth EstimationMonocular 3D Object Detection+2Categorical Depth Distribution Network for Monocular 3D Object Detection
Monocular 3D object detection is a key problem for autonomous vehicles, as it provides a solution with simple configuration compared to typical multi-sensor systems. The main challenge in monocular 3D detection lies in a…
3D Object DetectionAutonomous VehiclesDepth EstimationMonocular 3D Object Detection+3MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer
Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to assist 3D detection, but suffer from the…
3D Object Detection3D Object Detection From Monocular ImagesAutonomous DrivingMonocular 3D Object Detection+3ViewpointDepth: A New Dataset for Monocular Depth Estimation Under Viewpoint Shifts
Monocular depth estimation is a critical task for autonomous driving and many other computer vision applications. While significant progress has been made in this field, the effects of viewpoint shifts on depth estimatio…
Autonomous DrivingDepth EstimationHomography EstimationMonocular Depth Estimation+2