paper-with-me

홈 › Papers

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

2026-06-08 · Yimu Wang, Yee Man Choi, Barry Zhang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki arxiv

Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a model relied on the correct visual evidence. This gap is particularly important in multi-view driving scenes used for autonomous driving, where a model can produce a plausible answer while grounding it in the wrong camera view. We introduce a multi-view visual question answering benchmark for evaluating evidence-source identification: given six synchronized NuScenes views and a question, the model must identify the supporting camera view and answer the question. The benchmark contains 122 conflict-centric question-answer pairs from 73 scenes, spanning causality, counterfactual reasoning, and intent prediction. View labels are proposed by an automatic conflict-mining pipeline and manually verified by annotators. We evaluate three settings: camera-view selection, oracle QA given the golden view, and joint prediction in which the model selects a view and answers in one pass. Answers are evaluated in both multiple-choice and free-form formats, using exact match for structured predictions and an LLM judge for free-form responses. By explicitly separating visual-source identification from answer correctness, the benchmark exposes grounding failures that answer-only evaluation misses.

📄 PDF Abstract BibTeX arXiv:2606.09644

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringAutonomous DrivingVisual Reasoning

Similar Papers 제목 키워드 기반

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation

2024-12-16 · Debarshi Kundu

Consider the problem: ``If one man and one woman can produce one child in one year, how many children will be produced by one woman and three men in 0.5 years?" Current large language models (LLMs) such as GPT-4o, GPT-o1…

BenchmarkingDataset Generation

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

2026-07-01 · Jihyeok Jung, Jeewu Lee, Sanghyeop Kim, Chanhee Han 외 arxiv

Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-…

Scene Understanding

What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs

2025-05-15 · Xinlan Yan, Di wu, Yibin Lei, Christof Monz 외

In this paper, we introduce S-MedQA, an English medical question-answering (QA) dataset for benchmarking large language models in fine-grained clinical specialties. We use S-MedQA to check the applicability of a popular …

AllBenchmarkingMedical Question AnsweringMedQA+1

Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning

2024-06-11 · Shuvendu Roy, Yasaman Parhizkar, Franklin Ogidi, Vahid Reza Khazaie 외

We perform a comprehensive benchmarking of contrastive frameworks for learning multimodal representations in the medical domain. Through this study, we aim to answer the following research questions: (i) How transferable…

BenchmarkingContrastive LearningImage RetrievalImage to text+4

Less is More: Rejecting Unreliable Reviews for Product Question Answering

2020-07-09 · Shiwei Zhang, Xiuzhen Zhang, Jey Han Lau, Jeffrey Chan 외

Promptly and accurately answering questions on products is important for e-commerce applications. Manually answering product questions (e.g. on community question answering platforms) results in slow response and does no…

Community Question AnsweringConformal PredictionQuestion AnsweringRetrieval