Object-Aware Centroid Voting for Monocular 3D Object Detection
Monocular 3D object detection aims to detect objects in a 3D physical world from a single camera. However, recent approaches either rely on expensive LiDAR devices, or resort to dense pixel-wise depth estimation that causes prohibitive computational cost. In this paper, we propose an end-to-end trainable monocular 3D object detector without learning the dense depth. Specifically, the grid coordinates of a 2D box are first projected back to 3D space with the pinhole model as 3D centroids proposals. Then, a novel object-aware voting approach is introduced, which considers both the region-wise appearance attention and the geometric projection distribution, to vote the 3D centroid proposals for 3D object localization. With the late fusion and the predicted 3D orientation and dimension, the 3D bounding boxes of objects can be detected from a single RGB image. The method is straightforward yet significantly superior to other monocular-based methods. Extensive experimental results on the challenging KITTI benchmark validate the effectiveness of the proposed method.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionDepth EstimationMonocular 3D Object DetectionObjectobject-detectionObject DetectionObject LocalizationSimilar Papers 제목 키워드 기반
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
Reconstructing hand-held objects from a single RGB image is a challenging task in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, we employ a point cloud denoising diffusion mod…
DenoisingObjectObject ReconstructionNeighbor-Vote: Improving Monocular 3D Object Detection through Neighbor Distance Voting
As cameras are increasingly deployed in new application domains such as autonomous driving, performing 3D object detection on monocular images becomes an important task for visual scene understanding. Recent advances on …
3D Object DetectionAutonomous DrivingDepth EstimationMonocular 3D Object Detection+5Pixel Consensus Voting for Panoptic Segmentation
The core of our approach, Pixel Consensus Voting, is a framework for instance segmentation based on the Generalized Hough transform. Pixels cast discretized, probabilistic votes for the likely regions that contain instan…
Instance SegmentationPanoptic SegmentationSegmentationSemantic SegmentationMLCVNet: Multi-Level Context VoteNet for 3D Object Detection
In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize…
3D Object DetectionObjectobject-detectionObject DetectionDeep Hough Voting for 3D Object Detection in Point Clouds
Current 3D object detection methods are heavily influenced by 2D detectors. In order to leverage architectures in 2D detectors, they often convert 3D point clouds to regular grids (i.e., to voxel grids or to bird's eye v…
3D Object Detection3D Object Detection From Monocular ImagesObjectobject-detection+1