paper-with-me

Papers

Unsupervised 3D Perception with 2D Vision-Language Distillation for Autonomous Driving

2023-09-25 · ICCV 2023 1 · Mahyar Najibi, Jingwei Ji, Yin Zhou, Charles R. Qi, Xinchen Yan, Scott Ettinger, Dragomir Anguelov

Closed-set 3D perception models trained on only a pre-defined set of object categories can be inadequate for safety critical applications such as autonomous driving where new object types can be encountered after deployment. In this paper, we present a multi-modal auto labeling pipeline capable of generating amodal 3D bounding boxes and tracklets for training models on open-set categories without 3D human labels. Our pipeline exploits motion cues inherent in point cloud sequences in combination with the freely available 2D image-text pairs to identify and track all traffic participants. Compared to the recent studies in this domain, which can only provide class-agnostic auto labels limited to moving objects, our method can handle both static and moving objects in the unsupervised manner and is able to output open-vocabulary semantic labels thanks to the proposed vision-language knowledge distillation. Experiments on the Waymo Open Dataset show that our approach outperforms the prior work by significant margins on various unsupervised 3D perception tasks.

📄 PDF Abstract BibTeX arXiv:2309.14491

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingKnowledge Distillation

Similar Papers 제목 키워드 기반

RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillation

2026-06-12 · Xiangyu Huang, Zhenlin Hua, Han Zhou, Shounak Sural 외 arxiv

Vision-Language-Action (VLA) models have shown strong potential for end-to-end autonomous driving by jointly modeling visual perception, language reasoning, explainability and action prediction. However, their large visi…

Knowledge DistillationAutonomous Driving

ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving

2025-05-21 · Yunsheng Ma, Burhaneddin Yaman, Xin Ye, Mahmut Yurt 외

Recent advances have explored integrating large language models (LLMs) into end-to-end autonomous driving systems to enhance generalization and interpretability. However, most existing approaches are limited to either dr…

Autonomous Drivingcross-modal alignment

Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion

2026-05-08 · Yijing Wang, Ruonan Li, Qilin Wang, Rongqiang Zhao 외 arxiv

LiDAR-based semantic segmentation is essential for autonomous-driving perception, yet dense point-wise annotations are costly, and long-tailed outdoor scenes make small safety-critical objects difficult to discover witho…

Scene UnderstandingAutonomous DrivingPoint Clouds

VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion

2025-03-08 · Meng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin 외

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity cau…

3D Semantic Scene CompletionAutonomous DrivingLanguage ModelingLanguage Modelling+1

EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation

2026-03-10 · Jiajun Cao, Xiaoan Zhang, Xiaobao Wei, Liyuqiu Huang 외 arxiv

Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long-term planning.…

Autonomous Driving