paper-with-me

Papers

3D Dual-Fusion: Dual-Domain Dual-Query Camera-LiDAR Fusion for 3D Object Detection

2022-11-24 · Yecheol Kim, Konyul Park, Minwook Kim, Dongsuk Kum, Jun Won Choi

Fusing data from cameras and LiDAR sensors is an essential technique to achieve robust 3D object detection. One key challenge in camera-LiDAR fusion involves mitigating the large domain gap between the two sensors in terms of coordinates and data distribution when fusing their features. In this paper, we propose a novel camera-LiDAR fusion architecture called, 3D Dual-Fusion, which is designed to mitigate the gap between the feature representations of camera and LiDAR data. The proposed method fuses the features of the camera-view and 3D voxel-view domain and models their interactions through deformable attention. We redesign the transformer fusion encoder to aggregate the information from the two domains. Two major changes include 1) dual query-based deformable attention to fuse the dual-domain features interactively and 2) 3D local self-attention to encode the voxel-domain queries prior to dual-query decoding. The results of an experimental evaluation show that the proposed camera-LiDAR fusion architecture achieved competitive performance on the KITTI and nuScenes datasets, with state-of-the-art performances in some 3D object detection benchmarks categories.

📄 PDF Abstract BibTeX arXiv:2211.13529

Code (1)

rasd3/3D-Dual-Fusion 공식 구현 pytorch

Tasks

3D Object Detectionobject-detectionObject DetectionRobust 3D Object Detection

Similar Papers 제목 키워드 기반

DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition

2026-01-19 · Hanyu Zhu, Zhihao Zhan, Yuhang Ming, Liang Li 외 arxiv

One of the central challenges in visual place recognition (VPR) is learning a robust global representation that remains discriminative under large viewpoint changes, illumination variations, and severe domain shifts. Whi…

Visual Place Recognition

RestNet: Boosting Cross-Domain Few-Shot Segmentation with Residual Transformation Network

2023-08-25 · Xinyang Huang, Chuang Zhu, Wenkai Chen

Cross-domain few-shot segmentation (CD-FSS) aims to achieve semantic segmentation in previously unseen domains with a limited number of annotated samples. Although existing CD-FSS models focus on cross-domain feature tra…

Cross-Domain Few-ShotSemantic SegmentationTransfer Learning

Dual-Domain Homogeneous Fusion with Cross-Modal Mamba and Progressive Decoder for 3D Object Detection

2025-03-12 · Xuzhong Hu, Zaipeng Duan, Pei An, Jun Zhang 외

Fusing LiDAR point cloud features and image features in a homogeneous BEV space has been widely adopted for 3D object detection in autonomous driving. However, such methods are limited by the excessive compression of mul…

3D Object DetectionAutonomous DrivingDecoderFeature Compression+3

Dual Adversarial Alignment for Realistic Support-Query Shift Few-shot Learning

2023-09-05 · Siyang Jiang, Rui Fang, Hsi-Wen Chen, Wei Ding 외

Support-query shift few-shot learning aims to classify unseen examples (query set) to labeled data (support set) based on the learned embedding in a low-dimensional space under a distribution shift between the support se…

Few-Shot Learning

DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval

2025-05-23 · Yuxin Yang, Yinan Zhou, Yuxin Chen, Ziqi Zhang 외

Composed Image Retrieval (CIR) aims to retrieve target images from a gallery based on a reference image and modification text as a combined query. Recent approaches focus on balancing global information from two modaliti…

Image RetrievalRetrieval