Digging Into Output Representation for Monocular 3D Object Detection
Monocular 3D object detection aims to recognize and localize objects in 3D space from a single image. Recent researches have conducted remarkable advancements, while all of them follow a typical output representation in LiDAR-based 3D detection. However, in this paper, we argue that the existing discrete output representation is not suitable for monocular 3D detection. Specifically, monocular 3D detection has only two-dimensional information input while is required to output three-dimensional detections. This characteristic indicates that monocular 3D detection is inherently different from other typical detection tasks that have the same dimensional input and output. The dimension gap causes a large lower bound for the error of estimated depth. Therefore, we propose to reformulate the existing discrete output representation as a spatial probability distribution according to depth. This probability distribution considers the uncertainty caused by the absent depth dimension, allowing us to accurately and comprehensively represent objects in 3D space. Extensive experiments exhibit the superiority of our output representation. As a result, we have applied our method to 12 SOTA monocular 3D detectors, consistently boosting their average precision (AP) by ~ 20% relative improvements. The source code will be publicly available soon.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionMonocular 3D Object DetectionObjectobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, particularly due to its potential to enhance t…
3D geometryAutonomous VehiclesPredictionDeep Digging into the Generalization of Self-Supervised Monocular Depth Estimation
Self-supervised monocular depth estimation has been widely studied recently. Most of the work has focused on improving performance on benchmark datasets, such as KITTI, but has offered a few experiments on generalization…
Depth EstimationMonocular Depth EstimationDigging Into Self-Supervised Monocular Depth Estimation
Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation…
Camera Pose EstimationDepth EstimationImage ReconstructionMonocular Depth Estimation+4Digging Into Self-Supervised Monocular Depth Estimation
Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation…
Depth EstimationMonocular Depth EstimationSelf-Supervised LearningAlgorithmic Consequences of Particle Filters for Sentence Processing: Amplified Garden-Paths and Digging-In Effects
Under surprisal theory, linguistic representations affect processing difficulty only through the bottleneck of surprisal. Our best estimates of surprisal come from large language models, which have no explicit representa…