paper-with-me

홈 › Papers

V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views

2026-04-03 · Junwei You, Pei Li, Zhuoyu Jiang, Weizhe Tang, Zilin Huang, Rui Gan, Jiaxi Liu, Yan Zhao, Sikai Chen, Bin Ran arxiv

Multimodal large language models (MLLMs) have shown strong potential for autonomous driving, yet existing benchmarks remain largely ego-centric and therefore cannot systematically assess model performance in infrastructure-centric and cooperative driving conditions. In this work, we introduce V2X-QA, a real-world dataset and benchmark for evaluating MLLMs across vehicle-side, infrastructure-side, and cooperative viewpoints. V2X-QA is built around a view-decoupled evaluation protocol that enables controlled comparison under vehicle-only, infrastructure-only, and cooperative driving conditions within a unified multiple-choice question answering (MCQA) framework. The benchmark is organized into a twelve-task taxonomy spanning perception, prediction, and reasoning and planning, and is constructed through expert-verified MCQA annotation to enable fine-grained diagnosis of viewpoint-dependent capabilities. Benchmark results across ten representative state-of-the-art proprietary and open-source models show that viewpoint accessibility substantially affects performance, and infrastructure-side reasoning supports meaningful macroscopic traffic understanding. Results also indicate that cooperative reasoning remains challenging since it requires cross-view alignment and evidence integration rather than simply additional visual input. To address these challenges, we introduce V2X-MoE, a benchmark-aligned baseline with explicit view routing and viewpoint-specific LoRA experts. The strong performance of V2X-MoE further suggests that explicit viewpoint specialization is a promising direction for multi-view reasoning in autonomous driving. Overall, V2X-QA provides a foundation for studying multi-perspective reasoning, reliability, and cooperative physical intelligence in connected autonomous driving. The dataset and V2X-MoE resources are publicly available at: https://github.com/junwei0001/V2X-QA.

📄 PDF Abstract BibTeX arXiv:2604.02710

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingQuestion Answering

Similar Papers 제목 키워드 기반

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

2025-03-13 · Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang 외

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge.…

Multimodal Reasoning

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

2025-05-22 · Sara Ghaboura, Ketan More, Wafa Alghallabi, Omkar Thawakar 외

As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks remain focused on English, overlooking lang…

document understandingMultimodal ReasoningOptical Character Recognition (OCR)Visual Reasoning

MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning

2026-03-02 · Jiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang 외 arxiv

Recent progress in the reasoning capabilities of multimodal large language models (MLLMs) has empowered them to address more complex tasks such as scientific analysis and mathematical reasoning. Despite their promise, ML…

Mathematical ReasoningMultimodal Reasoning

FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging

2025-08-06 · Zichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang 외 arxiv

We present FinMMR, a novel bilingual multimodal benchmark tailored to evaluate the reasoning capabilities of multimodal large language models (MLLMs) in financial numerical reasoning tasks. Compared to existing benchmark…

MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification

2025-03-16 · Zhaopan Xu, Pengfei Zhou, Jiaxin Ai, Wangbo Zhao 외

Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, where the identification of process errors is vital for improving this ability. Recently, process-level reward models (PRMs) w…

Multimodal Reasoning