paper-with-me

홈 › Papers

Progressive Multi-Modal Fusion for Robust 3D Object Detection

2024-10-09 · Rohit Mohan, Daniele Cattaneo, Florian Drews, Abhinav Valada

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from both modalities either in Bird's Eye View (BEV) or Perspective View (PV), thus sacrificing complementary information such as height or geometric proportions. To address this limitation, we propose ProFusion3D, a progressive fusion framework that combines features in both BEV and PV at both intermediate and object query levels. Our architecture hierarchically fuses local and global features, enhancing the robustness of 3D object detection. Additionally, we introduce a self-supervised mask modeling pre-training strategy to improve multi-modal representation learning and data efficiency through three novel objectives. Extensive experiments on nuScenes and Argoverse2 datasets conclusively demonstrate the efficacy of ProFusion3D. Moreover, ProFusion3D is robust to sensor failure, demonstrating strong performance when only one modality is available.

📄 PDF Abstract BibTeX arXiv:2410.07475

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous DrivingObjectobject-detectionObject DetectionRepresentation LearningRobust 3D Object DetectionSensor Fusion

Similar Papers 제목 키워드 기반

Progressively Complementarity-Aware Fusion Network for RGB-D Salient Object Detection

2018-06-01 · CVPR 2018 6 · Hao Chen, Youfu Li

How to incorporate cross-modal complementarity sufficiently is the cornerstone question for RGB-D salient object detection. Previous works mainly address this issue by simply concatenating multi-modal features or combini…

object-detectionObject DetectionRGB-D Salient Object DetectionRGB Salient Object Detection+1

DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection

2024-09-24 · Jiaxin Ye, Junping Zhang, Hongming Shan

Depression is a common mental disorder that affects millions of people worldwide. Although promising, current multimodal methods hinge on aligned or aggregated multimodal fusion, suffering two significant limitations: (i…

Depression DetectionMamba

Depth-Cooperated Trimodal Network for Video Salient Object Detection

2022-02-12 · Yukang Lu, Dingyao Min, Keren Fu, Qijun Zhao

Depth can provide useful geographical cues for salient object detection (SOD), and has been proven helpful in recent RGB-D SOD methods. However, existing video salient object detection (VSOD) methods only utilize spatiot…

Objectobject-detectionObject DetectionOptical Flow Estimation+2

Dual-Domain Homogeneous Fusion with Cross-Modal Mamba and Progressive Decoder for 3D Object Detection

2025-03-12 · Xuzhong Hu, Zaipeng Duan, Pei An, Jun Zhang 외

Fusing LiDAR point cloud features and image features in a homogeneous BEV space has been widely adopted for 3D object detection in autonomous driving. However, such methods are limited by the excessive compression of mul…

3D Object DetectionAutonomous DrivingDecoderFeature Compression+3

Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection

2025-03-10 · Wentao Wu, Chenglong Li, Xiao Wang, Bin Luo 외

Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments, limiting detection performance…

Language ModelingLanguage ModellingLarge Language ModelObject+2