paper-with-me

홈 › Papers

V3LMA: Visual 3D-enhanced Language Model for Autonomous Driving

2025-04-30 · Jannik Lübberstedt, Esteban Rivera, Nico Uhlemann, Markus Lienkamp

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts their effectiveness in achieving a complete and safe understanding of dynamic surroundings. To address this, we introduce V3LMA, a novel approach that enhances 3D scene understanding by integrating Large Language Models (LLMs) with LVLMs. V3LMA leverages textual descriptions generated from object detections and video inputs, significantly boosting performance without requiring fine-tuning. Through a dedicated preprocessing pipeline that extracts 3D object data, our method improves situational awareness and decision-making in complex traffic scenarios, achieving a score of 0.56 on the LingoQA benchmark. We further explore different fusion strategies and token combinations with the goal of advancing the interpretation of traffic scenes, ultimately enabling safer autonomous driving systems.

📄 PDF Abstract BibTeX arXiv:2505.00156

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingDecision MakingLanguage ModelingLanguage ModellingScene Understanding

Similar Papers 제목 키워드 기반

Talk2BEV: Language-enhanced Bird's-eye View Maps for Autonomous Driving

2023-10-03 · Tushar Choudhary, Vikrant Dewangan, Shivam Chandhok, Shubham Priyadarshan 외

Talk2BEV is a large vision-language model (LVLM) interface for bird's-eye view (BEV) maps in autonomous driving contexts. While existing perception systems for autonomous driving scenarios have largely focused on a pre-d…

Autonomous DrivingDecision MakingLanguage ModelingLanguage Modelling+3

Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent

2024-11-08 · Linfeng He, Yiming Sun, Sihao Wu, Jiaxu Liu 외

In this paper, we propose a novel framework for enhancing visual comprehension in autonomous driving systems by integrating visual language models (VLMs) with additional visual perception module specialised in object det…

Autonomous DrivingLanguage ModelingLanguage Modellingobject-detection+3

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

2025-11-09 · Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang 외 arxiv

Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLMs). However, current LLM-based approache…

Autonomous Driving

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

2026-05-21 · Xiaodong Mei, Diankun Zhang, Hongwei Xie, Guang Chen 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene un…

Representation LearningImage ReconstructionScene UnderstandingAutonomous Driving

An interactive enhanced driving dataset for autonomous driving

2026-02-24 · Haojie Feng, Peizhi Zhang, Mengjie Tian, Xinrui Zhang 외 arxiv

The evolution of autonomous driving towards full automation demands robust interactive capabilities; however, the development of Vision-Language-Action (VLA) models is constrained by the sparsity of interactive scenarios…

Autonomous Driving