paper-with-me

홈 › Papers

V2X-REALM: Vision-Language Model-Based Robust End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling

2025-06-26 · Junwei You, Pei Li, Zhuoyu Jiang, Zilin Huang, Rui Gan, Haotian Shi, Bin Ran

Ensuring robust planning and decision-making under rare, diverse, and visually degraded long-tail scenarios remains a fundamental challenge for autonomous driving in urban environments. This issue becomes more critical in cooperative settings, where vehicles and infrastructure jointly perceive and reason across complex environments. To address this challenge, we propose V2X-REALM, a vision-language model (VLM)-based framework with adaptive multimodal learning for robust cooperative autonomous driving under long-tail scenarios. V2X-REALM introduces three core innovations: (i) a prompt-driven long-tail scenario generation and evaluation pipeline that leverages foundation models to synthesize realistic long-tail conditions such as snow and fog across vehicle- and infrastructure-side views, enriching training diversity efficiently; (ii) a gated multi-scenario adaptive attention module that modulates the visual stream using scenario priors to recalibrate ambiguous or corrupted features; and (iii) a multi-task scenario-aware contrastive learning objective that improves multimodal alignment and promotes cross-scenario feature separability. Extensive experiments demonstrate that V2X-REALM significantly outperforms existing baselines in robustness, semantic reasoning, safety, and planning accuracy under complex, challenging driving conditions, advancing the scalability of end-to-end cooperative autonomous driving.

📄 PDF Abstract BibTeX arXiv:2506.21041

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingContrastive LearningLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

2026-08-21 · Yitao Xu, Tong Wu, Yiyan Wu, Guoji Xu 외 arxiv

Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and la…

Trajectory PlanningAutonomous Driving

V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models

2024-08-17 · Junwei You, Haotian Shi, Zhuoyu Jiang, Zilin Huang 외

Vehicle-to-everything (V2X) cooperation has emerged as a promising paradigm to overcome the perception limitations of classical autonomous driving by leveraging information from both ego-vehicle and infrastructure sensor…

Autonomous DrivingContrastive LearningDecision MakingKnowledge Distillation+1

V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts

2025-09-22 · Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Yu-Chiang Frank Wang 외 arxiv

Current state-of-the-art autonomous vehicles could face safety-critical situations when their local sensors are occluded by large nearby objects on the road. Vehicle-to-vehicle (V2V) cooperative autonomous driving has be…

Autonomous VehiclesAutonomous Driving

V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multi-Modal Large Language Models

2025-02-14 · Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Stephen F. Smith 외

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To a…

Autonomous DrivingAutonomous VehiclesLarge Language ModelQuestion Answering

M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark

2025-05-10 · Morui Zhu, Yongqi Zhu, Yihao Zhu, Qi Chen 외

We introduce M$^3$CAD, a novel benchmark designed to advance research in generic cooperative autonomous driving. M$^3$CAD comprises 204 sequences with 30k frames, spanning a diverse range of cooperative driving scenarios…

Autonomous DrivingMotion Forecastingobject-detectionObject Detection