paper-with-me

Papers

Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving

2024-05-08 · Lingdong Kong, Xiang Xu, Jiawei Ren, Wenwei Zhang, Liang Pan, Kai Chen, Wei Tsang Ooi, Ziwei Liu

Efficient data utilization is crucial for advancing 3D scene understanding in autonomous driving, where reliance on heavily human-annotated LiDAR point clouds challenges fully supervised methods. Addressing this, our study extends into semi-supervised learning for LiDAR semantic segmentation, leveraging the intrinsic spatial priors of driving scenes and multi-sensor complements to augment the efficacy of unlabeled datasets. We introduce LaserMix++, an evolved framework that integrates laser beam manipulations from disparate LiDAR scans and incorporates LiDAR-camera correspondences to further assist data-efficient learning. Our framework is tailored to enhance 3D scene consistency regularization by incorporating multi-modality, including 1) multi-modal LaserMix operation for fine-grained cross-sensor interactions; 2) camera-to-LiDAR feature distillation that enhances LiDAR feature learning; and 3) language-driven knowledge guidance generating auxiliary supervisions using open-vocabulary models. The versatility of LaserMix++ enables applications across LiDAR representations, establishing it as a universally applicable solution. Our framework is rigorously validated through theoretical analysis and extensive experiments on popular driving perception datasets. Results demonstrate that LaserMix++ markedly outperforms fully supervised alternatives, achieving comparable accuracy with five times fewer annotations and significantly improving the supervised-only baselines. This substantial advancement underscores the potential of semi-supervised approaches in reducing the reliance on extensive labeled data in LiDAR-based 3D scene understanding systems.

📄 PDF Abstract BibTeX arXiv:2405.05258

Code (1)

ldkong1205/LaserMix 공식 구현 pytorch

Tasks

Autonomous DrivingLIDAR Semantic SegmentationScene UnderstandingSemantic Segmentation

Similar Papers 제목 키워드 기반

MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion

2025-12-15 · Minghui Hou, Wei-Hsing Huang, Shaofeng Liang, Daizong Liu 외 arxiv

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-la…

Key Information ExtractionMultimodal ReasoningScene UnderstandingAutonomous Driving

MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer

2025-08-20 · Guile Wu, David Huang, Dongfeng Bai, Bingbing Liu arxiv

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack th…

Scene UnderstandingAutonomous DrivingVideo Generation

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding

2025-07-07 · Jing Liang, Kasun Weerakoon, Daeun Song, Senthurbavan Kirubaharan 외 arxiv

We present MOSU, a novel autonomous long-range navigation system that enhances global navigation for mobile robots through multimodal perception and on-road scene understanding. MOSU addresses the outdoor robot navigatio…

Semantic SegmentationScene UnderstandingRobot Navigation

M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving

2024-03-19 · Dongyang Xu, Haokun Li, Qingfan Wang, Ziying Song 외

End-to-end autonomous driving has witnessed remarkable progress. However, the extensive deployment of autonomous vehicles has yet to be realized, primarily due to 1) inefficient multi-modal environment perception: how to…

Autonomous DrivingAutonomous VehiclesScene Understanding

DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding

2026-03-11 · Mingzhe Tao, Ruiping Liu, Junwei Zheng, Yufan Chen 외 arxiv

Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leverag…

Visual Question AnsweringScene UnderstandingAutonomous VehiclesAutonomous Driving