paper-with-me

홈 › Papers

A Generalized Multi-Modal Fusion Detection Framework

2023-03-13 · Leichao Cui, Xiuxian Li, Min Meng, Xiaoyu Mo

LiDAR point clouds have become the most common data source in autonomous driving. However, due to the sparsity of point clouds, accurate and reliable detection cannot be achieved in specific scenarios. Because of their complementarity with point clouds, images are getting increasing attention. Although with some success, existing fusion methods either perform hard fusion or do not fuse in a direct manner. In this paper, we propose a generic 3D detection framework called MMFusion, using multi-modal features. The framework aims to achieve accurate fusion between LiDAR and images to improve 3D detection in complex scenes. Our framework consists of two separate streams: the LiDAR stream and the camera stream, which can be compatible with any single-modal feature extraction network. The Voxel Local Perception Module in the LiDAR stream enhances local feature representation, and then the Multi-modal Feature Fusion Module selectively combines feature output from different streams to achieve better fusion. Extensive experiments have shown that our framework not only outperforms existing benchmarks but also improves their detection, especially for detecting cyclists and pedestrians on KITTI benchmarks, with strong robustness and generalization capabilities. Hopefully, our work will stimulate more research into multi-modal fusion for autonomous driving tasks.

📄 PDF Abstract BibTeX arXiv:2303.07064

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition

2026-08-03 · Guandi Wang, Ming Li, Yunsen Xing, Junle Liu arxiv

Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusion methods mainly weigh between two modalities and unify them in a u…

Object Detection

MaxCorrMGNN: A Multi-Graph Neural Network Framework for Generalized Multimodal Fusion of Medical Data for Outcome Prediction

2023-07-13 · Niharika S. D'Souza, Hongzhi Wang, Andrea Giovannini, Antonio Foncubierta-Rodriguez 외

With the emergence of multimodal electronic health records, the evidence for an outcome may be captured across multiple modalities ranging from clinical to imaging and genomic data. Predicting outcomes effectively requir…

Graph Neural Network

Generalized Hadamard-Product Fusion Operators for Visual Question Answering

2018-03-26 · Brendan Duke, Graham W. Taylor

We propose a generalized class of multimodal fusion operators for the task of visual question answering (VQA). We identify generalizations of existing multimodal fusion operators based on the Hadamard product, and show t…

Neural Architecture SearchQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detection

2025-03-03 · CVPR 2025 1 · Boyong He, Yuxiang Ji, Qianwen Ye, Zhuoyue Tan 외

Domain generalization (DG) for object detection aims to enhance detectors' performance in unseen scenarios. This task remains challenging due to complex variations in real-world applications. Recently, diffusion models h…

Domain AdaptationDomain Generalizationobject-detectionObject Detection+3

From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning

2025-03-08 · Shuangzhi Li, Junlong Shen, Lei Ma, Xingyu Li

LiDAR-based 3D object detection datasets have been pivotal for autonomous driving, yet they cover a limited range of objects, restricting the model's generalization across diverse deployment environments. To address this…

3D Object DetectionAutonomous DrivingCross-Domain Few-Shotcross-domain few-shot learning+4