Language-Guided 3D Object Detection in Point Cloud for Autonomous Driving
This paper addresses the problem of 3D referring expression comprehension (REC) in autonomous driving scenario, which aims to ground a natural language to the targeted region in LiDAR point clouds. Previous approaches for REC usually focus on the 2D or 3D-indoor domain, which is not suitable for accurately predicting the location of the queried 3D region in an autonomous driving scene. In addition, the upper-bound limitation and the heavy computation cost motivate us to explore a better solution. In this work, we propose a new multi-modal visual grounding task, termed LiDAR Grounding. Then we devise a Multi-modal Single Shot Grounding (MSSG) approach with an effective token fusion strategy. It jointly learns the LiDAR-based object detector with the language features and predicts the targeted region directly from the detector without any post-processing. Moreover, the image feature can be flexibly integrated into our approach to provide rich texture and color information. The cross-modal learning enforces the detector to concentrate on important regions in the point cloud by considering the informative language expressions, thus leading to much better accuracy and efficiency. Extensive experiments on the Talk2Car dataset demonstrate the effectiveness of the proposed methods. Our work offers a deeper insight into the LiDAR-based grounding task and we expect it presents a promising direction for the autonomous driving community.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionAutonomous Drivingobject-detectionObject DetectionReferring ExpressionReferring Expression ComprehensionVisual GroundingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection
The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across…
3D Object DetectionAllobject-detectionObject DetectionCG-SSD: Corner Guided Single Stage 3D Object Detection from LiDAR Point Cloud
At present, the anchor-based or anchor-free models that use LiDAR point clouds for 3D object detection use the center assigner strategy to infer the 3D bounding boxes. However, in a real world scene, the LiDAR can only a…
3D Object DetectionObjectobject-detectionObject DetectionAnchor-free 3D Single Stage Detector with Mask-Guided Attention for Point Cloud
Most of the existing single-stage and two-stage 3D object detectors are anchor-based methods, while the efficient but challenging anchor-free single-stage 3D object detection is not well investigated. Recent studies on 2…
2D Object Detection3D Object DetectionObjectobject-detection+1SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
3D object detection is essential for autonomous driving. As an emerging sensor, 4D imaging radar offers advantages as low cost, long-range detection, and accurate velocity measurement, making it highly suitable for objec…
3D Object DetectionAutonomous DrivingPoint CloudsVision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection
Accurate 3D object detection in LiDAR point clouds is crucial for autonomous driving systems. To achieve state-of-the-art performance, the supervised training of detectors requires large amounts of human-annotated data, …
3D Object DetectionAutonomous DrivingObjectobject-detection+2