paper-with-me

Papers

Evaluating Variance in Visual Question Answering Benchmarks

2025-08-04 · Nikitha SR arxiv

Multimodal large language models (MLLMs) have emerged as powerful tools for visual question answering (VQA), enabling reasoning and contextual understanding across visual and textual modalities. Despite their advancements, the evaluation of MLLMs on VQA benchmarks often relies on point estimates, overlooking the significant variance in performance caused by factors such as stochastic model outputs, training seed sensitivity, and hyperparameter configurations. This paper critically examines these issues by analyzing variance across 14 widely used VQA benchmarks, covering diverse tasks such as visual reasoning, text understanding, and commonsense reasoning. We systematically study the impact of training seed, framework non-determinism, model scale, and extended instruction finetuning on performance variability. Additionally, we explore Cloze-style evaluation as an alternate assessment strategy, studying its effectiveness in reducing stochasticity and improving reliability across benchmarks. Our findings highlight the limitations of current evaluation practices and advocate for variance-aware methodologies to foster more robust and reliable development of MLLMs.

📄 PDF Abstract BibTeX arXiv:2508.02645

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringVisual Reasoning

Similar Papers 제목 키워드 기반

InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

2025-05-25 · Minzhi Lin, Tianchi Xie, Mengchen Liu, Yilin Ye 외

Understanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, exist…

Chart UnderstandingQuestion AnsweringVisual Question Answering

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

2024-06-27 · Shubhankar Singh, Purvi Chaurasia, Yerram Varun, Pranshu Pandya 외

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities …

Decision MakingLogical ReasoningQuestion AnsweringSpatial Reasoning+2

On the Flip Side: Identifying Counterexamples in Visual Question Answering

2018-06-03 · Gabriel Grand, Aron Szanto, Yoon Kim, Alexander Rush

Visual question answering (VQA) models respond to open-ended natural language questions about images. While VQA is an increasingly popular area of research, it is unclear to what extent current VQA architectures learn ke…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models

2026-09-15 · Rwiddhi Chakraborty, Yinong, Wang, Cheng Zhang 외 arxiv

Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also bee…

Visual Question Answering

CBench: Towards Better Evaluation of Question Answering Over Knowledge Graphs

2021-04-05 · Abdelghny Orogat, Isabelle Liu, Ahmed El-Rob

Recently, there has been an increase in the number of knowledge graphs that can be only queried by experts. However, describing questions using structured queries is not straightforward for non-expert users who need to h…

BenchmarkingKnowledge GraphsQuestion Answering