paper-with-me

홈 › Papers

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

2022-10-18 · Xin Li, Botian Shi, Yuenan Hou, Xingjiao Wu, Tianlong Ma, Yikang Li, Liang He

Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse the image features with the point cloud features that are projected onto the 2D image plane or combine the sparse point cloud with dense image pixels. These fusion approaches often suffer from severe information loss, thus causing sub-optimal performance. To address these problems, we construct the homogeneous structure between the point cloud and images to avoid projective information loss by transforming the camera features into the LiDAR 3D space. In this paper, we propose a homogeneous multi-modal feature fusion and interaction method (HMFI) for 3D object detection. Specifically, we first design an image voxel lifter module (IVLM) to lift 2D image features into the 3D space and generate homogeneous image voxel features. Then, we fuse the voxelized point cloud features with the image features from different regions by introducing the self-attention based query fusion mechanism (QFM). Next, we propose a voxel feature interaction module (VFIM) to enforce the consistency of semantic information from identical objects in the homogeneous point cloud and image voxel representations, which can provide object-level alignment guidance for cross-modal feature fusion and strengthen the discriminative ability in complex backgrounds. We conduct extensive experiments on the KITTI and Waymo Open Dataset, and the proposed HMFI achieves better performance compared with the state-of-the-art multi-modal methods. Particularly, for the 3D detection of cyclist on the KITTI benchmark, HMFI surpasses all the published algorithms by a large margin.

📄 PDF Abstract BibTeX arXiv:2210.09615

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous Drivingobject-detectionObject Detection

Similar Papers 제목 키워드 기반

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition

2024-12-07 · Feng Li, Jiusong Luo, Wanjun Xia

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modaliti…

DiversityEmotion RecognitionRepresentation LearningSpeech Emotion Recognition

FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection

2025-01-08 · Guoxin Zhang, Ziying Song, Lin Liu, Zhonghong Ou

Multimodal 3D object detection has garnered considerable interest in autonomous driving. However, multimodal detectors suffer from dimension mismatches that derive from fusing 3D points with 2D pixels coarsely, which lea…

3D Object DetectionAutonomous Drivingmultimodal interactionobject-detection+1

LEGO: Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion

2024-10-02 · Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon 외

In computer vision tasks, features often come from diverse representations, domains, and modalities, such as text, images, and videos. Effectively fusing these features is essential for robust performance, especially wit…

Anomaly DetectionVideo Anomaly Detection

Enhancing Dyadic Relations with Homogeneous Graphs for Multimodal Recommendation

2023-01-28 · HongYu Zhou, Xin Zhou, Lingzi Zhang, Zhiqi Shen

User interaction data in recommender systems is a form of dyadic relation that reflects the preferences of users with items. Learning the representations of these two discrete sets of objects, users and items, is critica…

Graph LearningMultimodal RecommendationRecommendation Systems

Dual-Domain Homogeneous Fusion with Cross-Modal Mamba and Progressive Decoder for 3D Object Detection

2025-03-12 · Xuzhong Hu, Zaipeng Duan, Pei An, Jun Zhang 외

Fusing LiDAR point cloud features and image features in a homogeneous BEV space has been widely adopted for 3D object detection in autonomous driving. However, such methods are limited by the excessive compression of mul…

3D Object DetectionAutonomous DrivingDecoderFeature Compression+3