Cross-Modality 3D Object Detection
In this paper, we focus on exploring the fusion of images and point clouds for 3D object detection in view of the complementary nature of the two modalities, i.e., images possess more semantic information while point clouds specialize in distance sensing. To this end, we present a novel two-stage multi-modal fusion network for 3D object detection, taking both binocular images and raw point clouds as input. The whole architecture facilitates two-stage fusion. The first stage aims at producing 3D proposals through sparse point-wise feature fusion. Within the first stage, we further exploit a joint anchor mechanism that enables the network to utilize 2D-3D classification and regression simultaneously for better proposal generation. The second stage works on the 2D and 3D proposal regions and fuses their dense features. In addition, we propose to use pseudo LiDAR points from stereo matching as a data augmentation method to densify the LiDAR points, as we observe that objects missed by the detection network mostly have too few points especially for far-away objects. Our experiments on the KITTI dataset show that the proposed multi-stage fusion helps the network to learn better representations.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Classification3D Object DetectionData AugmentationObjectobject-detectionObject DetectionStereo MatchingSimilar Papers 제목 키워드 기반
DEYOLO: Dual-Feature-Enhancement YOLO for Cross-Modality Object Detection
Object detection in poor-illumination environments is a challenging task as objects are usually not clearly visible in RGB images. As infrared images provide additional clear edge information that complements RGB images,…
Objectobject-detectionObject DetectionCross-Modality Fusion Transformer for Multispectral Object Detection
Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effectiv…
Multispectral Object DetectionObjectobject-detectionObject Detection+1\emph{cm}SalGAN: RGB-D Salient Object Detection with Cross-View Generative Adversarial Networks
Image salient object detection (SOD) is an active research topic in computer vision and multimedia area. Fusing complementary information of RGB and depth has been demonstrated to be effective for image salient object de…
Edge DetectionGenerative Adversarial NetworkObjectobject-detection+5CFMW: Cross-modality Fusion Mamba for Multispectral Object Detection under Adverse Weather Conditions
Cross-modality images that integrate visible-infrared spectra cues can provide richer complementary information for object detection. Despite this, existing visible-infrared object detection methods severely degrade in s…
MambaMultispectral Object DetectionObjectobject-detection+1Fusion-Mamba for Cross-modality Object Detection
Cross-modality fusing complementary information from different modalities effectively improves object detection performance, making it more useful and robust for a wider range of applications. Existing fusion strategies …
MambaObjectobject-detectionObject Detection