paper-with-me

Papers

MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection

2024-08-12 · Zitian Wang, Zehao Huang, Yulu Gao, Naiyan Wang, Si Liu

The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages--cameras provide rich texture information and LiDAR offers precise 3D spatial data--relying on a single modality often leads to performance limitations. This paper introduces MV2DFusion, a multi-modal detection framework that integrates the strengths of both worlds through an advanced query-based fusion mechanism. By introducing an image query generator to align with image-specific attributes and a point cloud query generator, MV2DFusion effectively combines modality-specific object semantics without biasing toward one single modality. Then the sparse fusion process can be accomplished based on the valuable object semantics, ensuring efficient and accurate object detection across various scenarios. Our framework's flexibility allows it to integrate with any image and point cloud-based detectors, showcasing its adaptability and potential for future advancements. Extensive evaluations on the nuScenes and Argoverse2 datasets demonstrate that MV2DFusion achieves state-of-the-art performance, particularly excelling in long-range detection scenarios.

📄 PDF Abstract BibTeX arXiv:2408.05945

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous VehiclesObjectobject-detectionObject DetectionRobust 3D Object Detection

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object Detection

2022-09-07 · CVPR 2023 1 · Yang Jiao, Zequn Jie, Shaoxiang Chen, Jingjing Chen 외

Fusing LiDAR and camera information is essential for achieving accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric a…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection

SPDFusion: An Infrared and Visible Image Fusion Network Based on a Non-Euclidean Representation of Riemannian Manifolds

2024-11-16 · Huan Kang, Hui Li, Tianyang Xu, Rui Wang 외

Euclidean representation learning methods have achieved commendable results in image fusion tasks, which can be attributed to their clear advantages in handling with linear space. However, data collected from a realistic…

Infrared And Visible Image FusionRepresentation Learning

MaskedFusion: Mask-based 6D Object Pose Estimation

2019-11-18 · Nuno Pereira, Luís A. Alexandre

MaskedFusion is a framework to estimate the 6D pose of objects using RGB-D data, with an architecture that leverages multiple sub-tasks in a pipeline to achieve accurate 6D poses. 6D pose estimation is an open challenge …

6D Pose Estimation6D Pose Estimation using RGBDObjectPose Estimation

CMDFusion: Bidirectional Fusion Network with Cross-modality Knowledge Distillation for LIDAR Semantic Segmentation

2023-07-09 · Jun Cen, Shiwei Zhang, Yixuan Pei, Kun Li 외

2D RGB images and 3D LIDAR point clouds provide complementary knowledge for the perception system of autonomous vehicles. Several 2D and 3D fusion methods have been explored for the LIDAR semantic segmentation task, but …

Autonomous VehiclesKnowledge DistillationLIDAR Semantic SegmentationSemantic Segmentation

BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering

2026-06-15 · Yi-Ruei Liu, Jie-Ying Lee, Zheng-Hui Huang, Yu-Lun Liu 외 arxiv

Inverse rendering of urban scenes from captured videos enables numerous applications, including content creation and autonomous driving simulation. Physically-based rendering methods follow and control lighting physics, …

Autonomous DrivingInverse Rendering