paper-with-me

홈 › Papers

GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving

2026-06-01 · Yingzi Ma, Chaowei Xiao, Ming Jiang arxiv

Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains underexplored, raising uncertainties about their deployment across diverse global settings. We therefore introduce GeoDrive-Bench, a novel benchmark that enables the systematic investigation of VLMs' geo-culturally grounded driving reasoning. We curated 5,053 human-validated multiple-choice QA pairs across six countries covering diverse driving cultures. Specifically, we emphasize four driving tasks: perception, prediction, planning, and region reasoning. Each question requires models to infer the correct driving behavior from visual evidence and local traffic conventions without explicit country labels. Beyond evaluation, we further design a distillation algorithm that injects region-specific traffic-rule knowledge into the internal representations of VLMs, enabling models to better align visual scene understanding with local driving policies. Experiments on nine state-of-the-art VLMs show substantial performance variations across geo-driving cultures for each task, while our proposed baseline models exhibit improved geo-cultural reasoning across regions. These results suggest that current VLMs still lack robust region-aware driving intelligence and highlight GeoDrive-Bench as a diagnostic and training-oriented testbed for deployable autonomous driving foundation models.

📄 PDF Abstract BibTeX arXiv:2606.02774

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningScene UnderstandingAutonomous Driving

Similar Papers 제목 키워드 기반

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

2025-05-30 · Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky 외

Cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meani…

BenchmarkingMachine TranslationMultimodal Machine TranslationTranslation

SAGE: An Expert-Annotated South Asian GI Endoscopy Dataset for Multimodal Learning and Hallucination Analysis

2026-06-20 · Niyoj Oli, Sachin Acharya, Sandesh Pokhrel, Sanjay Bhandari 외 arxiv

Gastrointestinal cancers represent a growing health burden in the South Asian region, driven largely by rapid changes in socio-economic conditions and lifestyle habits. However, early diagnosis remains limited by inadequ…

Multi-Label ClassificationMulti-class ClassificationVisual Question AnsweringImage Captioning

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

2026-08-03 · Xianjing Han, Yuhan Su, Yang Deng, Dong Ma 외 arxiv

Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, …

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

2025-05-28 · Anthony Chen, Wenzhao Zheng, Yida Wang, Xueyang Zhang 외

Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anti…

3D geometryAutonomous DrivingAutonomous NavigationOcclusion Handling

Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems

2025-09-25 · Junfeng Yan, Biao Wu, Meng Fang, Ling Chen arxiv

Multimodal agents have demonstrated strong performance in general GUI interactions, but their application in automotive systems has been largely unexplored. In-vehicle GUIs present distinct challenges: drivers' limited a…