Investigating Vision-Language Model for Point Cloud-based Vehicle Classification
Heavy-duty trucks pose significant safety challenges due to their large size and limited maneuverability compared to passenger vehicles. A deeper understanding of truck characteristics is essential for enhancing the safety perspective of cooperative autonomous driving. Traditional LiDAR-based truck classification methods rely on extensive manual annotations, which makes them labor-intensive and costly. The rapid advancement of large language models (LLMs) trained on massive datasets presents an opportunity to leverage their few-shot learning capabilities for truck classification. However, existing vision-language models (VLMs) are primarily trained on image datasets, which makes it challenging to directly process point cloud data. This study introduces a novel framework that integrates roadside LiDAR point cloud data with VLMs to facilitate efficient and accurate truck classification, which supports cooperative and safe driving environments. This study introduces three key innovations: (1) leveraging real-world LiDAR datasets for model development, (2) designing a preprocessing pipeline to adapt point cloud data for VLM input, including point cloud registration for dense 3D rendering and mathematical morphological techniques to enhance feature representation, and (3) utilizing in-context learning with few-shot prompting to enable vehicle classification with minimally labeled training data. Experimental results demonstrate encouraging performance of this method and present its potential to reduce annotation efforts while improving classification accuracy.
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous DrivingClassificationFew-Shot LearningIn-Context LearningLanguage ModelingLanguage ModellingPoint Cloud RegistrationSimilar Papers 제목 키워드 기반
Lateral Ego-Vehicle Control without Supervision using Point Clouds
Existing vision based supervised approaches to lateral vehicle control are capable of directly mapping RGB images to the appropriate steering commands. However, they are prone to suffering from inadequate robustness in r…
Visual OdometryVANETs Meet Autonomous Vehicles: A Multimodal 3D Environment Learning Approach
In this paper, we design a multimodal framework for object detection, recognition and mapping based on the fusion of stereo camera frames, point cloud Velodyne Lidar scans, and Vehicle-to-Vehicle (V2V) Basic Safety Messa…
Autonomous Vehiclesobject-detectionObject DetectionImage2Point: 3D Point-Cloud Understanding with 2D Image Pretrained Models
3D point-clouds and 2D images are different visual representations of the physical world. While human vision can understand both representations, computer vision models designed for 2D image and 3D point-cloud understand…
3D Point Cloud ClassificationPoint Cloud ClassificationScene SegmentationMulti-view Structural Convolution Network for Domain-Invariant Point Cloud Recognition of Autonomous Vehicles
Point cloud representation has recently become a research hotspot in the field of computer vision and has been utilized for autonomous vehicles. However, adapting deep learning networks for point cloud data recognition i…
Autonomous VehiclesDeconvolutional Networks for Point-Cloud Vehicle Detection and Tracking in Driving Scenarios
Vehicle detection and tracking is a core ingredient for developing autonomous driving applications in urban scenarios. Recent image-based Deep Learning (DL) techniques are obtaining breakthrough results in these percepti…
Autonomous DrivingAutonomous VehiclesMulti-Object TrackingObject Tracking+1