paper-with-me

Papers

KAN-RCBEVDepth: A multi-modal fusion algorithm in object detection for autonomous driving

2024-08-04 · Zhihao Lai, Chuanhao Liu, Shihui Sheng, Zhiqiang Zhang

Accurate 3D object detection in autonomous driving is critical yet challenging due to occlusions, varying object sizes, and complex urban environments. This paper introduces the KAN-RCBEVDepth method, an innovative approach aimed at enhancing 3D object detection by fusing multimodal sensor data from cameras, LiDAR, and millimeter-wave radar. Our unique Bird's Eye View-based approach significantly improves detection accuracy and efficiency by seamlessly integrating diverse sensor inputs, refining spatial relationship understanding, and optimizing computational procedures. Experimental results show that the proposed method outperforms existing techniques across multiple detection metrics, achieving a higher Mean Distance AP (0.389, 23\% improvement), a better ND Score (0.485, 17.1\% improvement), and a faster Evaluation Time (71.28s, 8\% faster). Additionally, the KAN-RCBEVDepth method significantly reduces errors compared to BEVDepth, with lower Transformation Error (0.6044, 13.8\% improvement), Scale Error (0.2780, 2.6\% improvement), Orientation Error (0.5830, 7.6\% improvement), Velocity Error (0.4244, 28.3\% improvement), and Attribute Error (0.2129, 3.2\% improvement). These findings suggest that our method offers enhanced accuracy, reliability, and efficiency, making it well-suited for dynamic and demanding autonomous driving scenarios. The code will be released in \url{https://github.com/laitiamo/RCBEVDepth-KAN}.

📄 PDF Abstract BibTeX arXiv:2408.02088

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAttributeAutonomous DrivingObjectobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Multi-Modal 3D Object Detection in Autonomous Driving: a Survey

2021-06-24 · Yingjie Wang, Qiuyu Mao, Hanqi Zhu, Jiajun Deng 외

In this survey, we first introduce the background of popular sensors used for self-driving, their data properties, and the corresponding object detection algorithms. Next, we discuss existing datasets that can be used fo…

3D Object DetectionAutonomous DrivingObjectobject-detection+3

MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection

2025-01-12 · Kaiying Yan, Moyang Liu, Yukun Liu, Ruibo Fu 외

Multimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal information lead to intensified optimization conf…

Fake News Detection

CFTrack: Center-based Radar and Camera Fusion for 3D Multi-Object Tracking

2021-07-11 · Ramin Nabati, Landon Harris, Hairong Qi

3D multi-object tracking is a crucial component in the perception system of autonomous driving vehicles. Tracking all dynamic objects around the vehicle is essential for tasks such as obstacle avoidance and path planning…

3D Multi-Object TrackingAutonomous DrivingAutonomous VehiclesMulti-Object Tracking+5

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

2026-08-26 · Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang 외 arxiv

Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, and sonar images limited by less object structural information. While…

Object Detection

E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection

2024-03-14 · Jiaqing Zhang, Mingxiang Cao, Weiying Xie, Jie Lei 외

Multimodal image fusion and object detection are crucial for autonomous driving. While current methods have advanced the fusion of texture details and semantic information, their complex training processes hinder broader…

Autonomous DrivingObjectobject-detectionObject Detection+1