paper-with-me

홈 › Papers

What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks

2026-07-11 · Guanhua Ye, Niu Jingbin, Yan Li, Meiyu Liang, Zhe Xue, Yingxia Shao, Yawen Li arxiv

Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision--language models and six benchmarks, using a human-validated semantic judge (97.6% precision) to audit over 37k official errors. A second text-only judge reproduces the same benchmark-level false-negative pattern, showing that the effect is not an artifact of a single audit model. On text-rich benchmarks, up to half of these errors are semantically acceptable answers penalized purely for surface-form mismatch. This instability is structured by answer type: extractive and multi-span answers are far more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further destabilize official outcomes, flipping item-level correctness at substantial rates without changing the underlying task. A deterministic CPU-only contract repair confirms that the undercount is partially recoverable. These findings imply that official short-answer VQA scores should be accompanied by semantic audits and answer-type diagnostics to remain interpretable.

📄 PDF Abstract BibTeX arXiv:2607.10240

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perceptual Score: What Data Modalities Does Your Model Perceive?

2021-10-27 · NeurIPS 2021 12 · Itai Gat, Idan Schwartz, Alexander Schwing

Machine learning advances in the last decade have relied significantly on large-scale datasets that continue to grow in size. Increasingly, those datasets also contain different data modalities. However, large multi-moda…

Question AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)

When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models

2026-04-24 · Pruthvinath Jeripity Venkata arxiv

When you ask an AI assistant for advice about your career, your marriage, or a conflict with your family, does it give you the same answer regardless of where you are from? We tested this systematically by presenting thr…

Your Answer is Incorrect... Would you like to know why? Introducing a Bilingual Short Answer Feedback Dataset

2022-05-01 · ACL 2022 5 · Anna Filighera, Siddharth Parihar, Tim Steuer, Tobias Meuser 외

Handing in a paper or exercise and merely receiving “bad” or “incorrect” as feedback is not very helpful when the goal is to improve. Unfortunately, this is currently the kind of feedback given by Automatic Short Answer …

automatic short answer grading

Your Answer is Incorrect... Would you like to know why? Introducing a Bilingual Short Answer Feedback Dataset

2021-12-17 · ACL ARR December 2022 12 · Anonymous

Handing in a paper or exercise and merely receiving a "bad" or "incorrect" as feedback is not very helpful when the goal is to improve. Unfortunately, this is currently the kind of feedback given by many Automatic Short …

automatic short answer grading

Verifying Tree Ensembles by Reasoning about Potential Instances

2020-01-31 · Laurens Devos, Wannes Meert, Jesse Davis

Imagine being able to ask questions to a black box model such as "Which adversarial examples exist?", "Does a specific attribute have a disproportionate effect on the model's prediction?" or "What kind of predictions cou…

AttributeFairness