paper-with-me

홈 › Papers

FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection

2023-09-11 · Chunyong Hu, Hang Zheng, Kun Li, Jianyun Xu, Weibo Mao, Maochun Luo, Lingxuan Wang, Mingxia Chen, Qihao Peng, Kaixuan Liu, Yiru Zhao, Peihan Hao, Minzhe Liu, Kaicheng Yu

Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features into the bird's eye view space and may lose certain information on Z-axis, thus leading to inferior performance. To this end, we propose a novel end-to-end multi-modal fusion transformer-based framework, dubbed FusionFormer, that incorporates deformable attention and residual structures within the fusion encoding module. Specifically, by developing a uniform sampling strategy, our method can easily sample from 2D image and 3D voxel features spontaneously, thus exploiting flexible adaptability and avoiding explicit transformation to the bird's eye view space during the feature concatenation process. We further implement a residual structure in our feature encoder to ensure the model's robustness in case of missing an input modality. Through extensive experiments on a popular autonomous driving benchmark dataset, nuScenes, our method achieves state-of-the-art single model performance of 72.6% mAP and 75.1% NDS in the 3D object detection task without test time augmentation.

📄 PDF Abstract BibTeX arXiv:2309.05257

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous DrivingDepth EstimationDepth Predictionobject-detectionObject Detection

Similar Papers 제목 키워드 기반

FusionFormer: Fusing Operations in Transformer for Efficient Streaming Speech Recognition

2022-10-31 · Xingchen Song, Di wu, BinBin Zhang, Zhiyong Wu 외

The recently proposed Conformer architecture which combines convolution with attention to capture both local and global dependencies has become the \textit{de facto} backbone model for Automatic Speech Recognition~(ASR).…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Multi-View 3D Object Detection Network for Autonomous Driving

2016-11-23 · CVPR 2017 7 · Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li 외

This paper aims at high-accuracy 3D object detection in autonomous driving scenario. We propose Multi-View 3D networks (MV3D), a sensory-fusion framework that takes both LIDAR point cloud and RGB images as input and pred…

3D Object DetectionAutonomous DrivingObjectobject-detection+2

IRFusionFormer: Enhancing Pavement Crack Segmentation with RGB-T Fusion and Topological-Based Loss

2024-09-30 · Ruiqiang Xiao, Xiaohu Chen

Crack segmentation is crucial in civil engineering, particularly for assessing pavement integrity and ensuring the durability of infrastructure. While deep learning has advanced RGB-based segmentation, performance degrad…

Crack SegmentationSegmentation

(Fusionformer):Exploiting the Joint Motion Synergy with Fusion Network Based On Transformer for 3D Human Pose Estimation

2022-10-08 · Xinwei Yu, Xiaohua Zhang

For the current 3D human pose estimation task, a group of methods mainly learn the rules of 2D-3D projection from spatial and temporal correlation. However, earlier methods model the global features of the entire body jo…

3D Human Pose EstimationPose Estimation

Fusing Bird View LIDAR Point Cloud and Front View Camera Image for Deep Object Detection

2017-11-17 · Zining Wang, Wei Zhan, Masayoshi Tomizuka

We propose a new method for fusing a LIDAR point cloud and camera-captured images in the deep convolutional neural network (CNN). The proposed method constructs a new layer called non-homogeneous pooling layer to transfo…

3D Object DetectionAutonomous DrivingObjectobject-detection+1