paper-with-me

홈 › Papers

VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion

2025-02-25 · Pei Liu, Haipeng Liu, Haichao Liu, Xin Liu, Jinxin Ni, Jun Ma

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D observations into 3D space. In this sense, it hinders their effective deployment in dynamic and complex environments. Leveraging the superior scene understanding and reasoning abilities of Vision-Language Models (VLMs), we propose VLM-E2E, a novel framework that uses the VLMs to enhance training by providing attentional cues. Our method integrates textual representations into Bird's-Eye-View (BEV) features for semantic supervision, which enables the model to learn richer feature representations that explicitly capture the driver's attentional semantics. By focusing on attentional semantics, VLM-E2E better aligns with human-like driving behavior, which is critical for navigating dynamic and complex environments. Furthermore, we introduce a BEV-Text learnable weighted fusion strategy to address the issue of modality importance imbalance in fusing multimodal information. This approach dynamically balances the contributions of BEV and text features, ensuring that the complementary information from visual and textual modality is effectively utilized. By explicitly addressing the imbalance in multimodal fusion, our method facilitates a more holistic and robust representation of driving environments. We evaluate VLM-E2E on the nuScenes dataset and demonstrate its superiority over state-of-the-art approaches, showcasing significant improvements in performance.

📄 PDF Abstract BibTeX arXiv:2502.18042

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingNavigateScene Understanding

Similar Papers 제목 키워드 기반

Driver Assistance System Based on Multimodal Data Hazard Detection

2025-02-05 · Long Zhouxiang, Ovanes Petrosian

Autonomous driving technology has advanced significantly, yet detecting driving anomalies remains a major challenge due to the long-tailed distribution of driving events. Existing methods primarily rely on single-modal r…

Autonomous Driving

SuperDriverAI: Towards Design and Implementation for End-to-End Learning-based Autonomous Driving

2023-05-14 · Shunsuke Aoki, Issei Yamamoto, Daiki Shiotsuka, Yuichi Inoue 외

Fully autonomous driving has been widely studied and is becoming increasingly feasible. However, such autonomous driving has yet to be achieved on public roads, because of various uncertainties due to surrounding human d…

Autonomous Driving

MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving

2026-02-25 · Lingjun Zhang, Yujian Yuan, Changjie Wu, Xinyuan Chang 외 arxiv

Vision-Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end-to-end autonomous driving systems. Chain-of-Thought (CoT), as VLM's widely used reasoning strategy, is facing critical challenge…

Multimodal ReasoningTrajectory PlanningAutonomous Driving

InVDriver: Intra-Instance Aware Vectorized Query-Based Autonomous Driving Transformer

2025-02-25 · Bo Zhang, Heye Huang, Chunyang Liu, Yaqin Zhang 외

End-to-end autonomous driving with its holistic optimization capabilities, has gained increasing traction in academia and industry. Vectorized representations, which preserve instance-level topological information while …

Autonomous DrivingComputational Efficiency

DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance

2025-12-16 · Shreedhar Govil, Didier Stricker, Jason Rambach arxiv

Predicting driver attention is a critical problem for developing explainable autonomous driving systems and understanding driver behavior in mixed human-autonomous vehicle traffic scenarios. Although significant progress…

Semantic SegmentationAutonomous Driving