paper-with-me

홈 › Papers

Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models

2026-05-20 · Steven Chen, Shivesh Khaitan, Nemanja Djuric arxiv

We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicle Make and Model Recognition (VMMR) methods. The proposed approach utilizes a Vision Language Model (VLM) to both infer a vehicle's make, model, and generation from image crops, and output accurate 3D bounding box dimensions to seed manual labeling. We evaluate the impact of iterative prompt engineering and the choice of different VLMs on both vehicle bounding box inference and make/model/generation recognition. When compared to strong baselines, the proposed approach not only shows high accuracy, but also excels in mitigating specific failure modes where VLMs provide better dimensions than initial lidar-aided human annotated labels (e.g., in cases of significant vehicle occlusion). Experiments on both public and proprietary data strongly suggest that our conclusions are generalizable across different labelers and datasets. The results demonstrate that integrating VLMs into the labeling process can reduce manual labeling time while increasing label quality.

📄 PDF Abstract BibTeX arXiv:2605.21747

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Detection of Active Emergency Vehicles using Per-Frame CNNs and Output Smoothing

2022-12-28 · Meng Fan, Craig Bidstrup, Zhaoen Su, Jason Owens 외

While inferring common actor states (such as position or velocity) is an important and well-explored task of the perception system aboard a self-driving vehicle (SDV), it may not always provide sufficient information to …

Data AugmentationPosition

A Method for Vehicle Collision Risk Assessment through Inferring Driver's Braking Actions in Near-Crash Situations

2020-04-28

Driving information and data under potential vehicle crashes create opportunities for extensive real-world observations of driver behaviors and relevant factors that significantly influence the driving safety in emergenc…

AttributeCollision Avoidance

Inferring Driving Maps by Deep Learning-based Trail Map Extraction

2025-05-15 · Michael Hubbertz, Pascal Colling, Qi Han, Tobias Meisen

High-definition (HD) maps offer extensive and accurate environmental information about the driving scene, making them a crucial and essential element for planning within autonomous driving systems. To avoid extensive eff…

Autonomous DrivingDeep Learning

Uncertainty-Aware Vehicle Orientation Estimation for Joint Detection-Prediction Models

2020-11-05 · Henggang Cui, Fang-Chieh Chou, Jake Charland, Carlos Vallespi-Gonzalez 외

Object detection is a critical component of a self-driving system, tasked with inferring the current states of the surrounding traffic actors. While there exist a number of studies on the problem of inferring the positio…

motion predictionobject-detectionObject DetectionPrediction

Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds

2024-09-10 · Mu Cai, Chenxu Luo, Yong Jae Lee, Xiaodong Yang

3D perception in LiDAR point clouds is crucial for a self-driving vehicle to properly act in 3D environment. However, manually labeling point clouds is hard and costly. There has been a growing interest in self-supervise…

3D Object Detection3D Semantic SegmentationAutonomous DrivingContrastive Learning+4