paper-with-me

홈 › Papers

Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models

2025-09-20 · Md. Atabuzzaman, Ali Asgarov, Chris Thomas arxiv

Large Vision-Language Models (LVLMs) have achieved strong performance on vision-language tasks, particularly Visual Question Answering (VQA). While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor specific option tokens (e.g., "A") or positions, remains underexplored. In this paper, we investigate both the presence and nature of selection bias in LVLMs through fine-grained MCQA benchmarks spanning easy, medium, and hard difficulty levels, defined by the semantic similarity of the options. We further propose an inference-time logit-level debiasing method that estimates an ensemble bias vector from general and contextual prompts and applies confidence-adaptive corrections to the model's output. Our method mitigates bias without retraining and is compatible with frozen LVLMs. Extensive experiments across several state-of-the-art models reveal consistent selection biases that intensify with task difficulty, and show that our mitigation approach significantly reduces bias while improving accuracy in challenging settings. This work offers new insights into the limitations of LVLMs in MCQA and presents a practical approach to improve their robustness in fine-grained visual reasoning. Datasets and code are available at: https://github.com/Atabuzzaman/Selection-Bias-of-LVLMs

📄 PDF Abstract BibTeX arXiv:2509.16805

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringSemantic SimilarityVisual Reasoning

Similar Papers 제목 키워드 기반

Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

2024-10-18 · Olga Loginova, Oleksandr Bezrukov, Alexey Kravets

Evaluating Video Language Models (VLMs) is a challenging task. Due to its transparency, Multiple-Choice Question Answering (MCQA) is widely used to measure the performance of these models through accuracy. However, exist…

FairnessMultiple-choiceMultiple Choice Question Answering (MCQA)Question Answering+1

Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs

2025-09-24 · Shree Harsha Bokkahalli Satish, Gustav Eje Henter, Éva Székely arxiv

Recent work in benchmarking bias and fairness in speech large language models (SpeechLLMs) has relied heavily on multiple-choice question answering (MCQA) formats. The model is tasked to choose between stereotypical, ant…

Question Answering

Afri-MCQA: Multimodal Cultural Question Answering for African Languages

2026-01-09 · Atnafu Lambebo Tonja, Srija Anand, Emilio Villa-Cueva, Israel Abebe Azime 외 arxiv

Africa is home to over one-third of the world's languages, yet remains underrepresented in AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering benchmark covering 7.5k Q&A pairs across …

Question Answering

Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?

2024-07-02 · Nishant Balepur, Rachel Rudinger

Recent work shows that large language models (LLMs) can answer multiple-choice questions using only the choices, but does this mean that MCQA leaderboard rankings of LLMs are largely influenced by abilities in choices-on…

Graph MiningLanguage ModelingLanguage ModellingLarge Language Model+1

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

2026-06-10 · Antoni Lasik, Jakub Pokrywka, Łukasz Grzybowski, Jeremi Ignacy Kaczmarek 외 arxiv

Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical ability due to guessing strategies and answer biases. To address these l…

Question Answering