paper-with-me

Papers

Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning

2024-10-06 · Shramay Palta, Nishant Balepur, Peter Rankel, Sarah Wiegreffe, Marine Carpuat, Rachel Rudinger

Questions involving commonsense reasoning about everyday situations often admit many $\textit{possible}$ or $\textit{plausible}$ answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning require a hard selection of a single correct answer, which, in principle, should represent the $\textit{most}$ plausible answer choice. On $250$ MCQ items sampled from two commonsense reasoning benchmarks, we collect $5,000$ independent plausibility judgments on answer choices. We find that for over 20% of the sampled MCQs, the answer choice rated most plausible does not match the benchmark gold answers; upon manual inspection, we confirm that this subset exhibits higher rates of problems like ambiguity or semantic mismatch between question and answer choices. Experiments with LLMs reveal low accuracy and high variation in performance on the subset, suggesting our plausibility criterion may be helpful in identifying more reliable benchmark items for commonsense evaluation.

📄 PDF Abstract BibTeX arXiv:2410.10854

Code (1)

shramay-palta/commonsense-mcq-plausibility 공식 구현

Tasks

Multiple-choice

Similar Papers 제목 키워드 기반

Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers

2025-10-09 · Nishant Balepur, Atrey Desai, Rachel Rudinger arxiv

Large language models (LLMs) now give reasoning before answering, excelling in tasks like multiple-choice question answering (MCQA). Yet, a concern is that LLMs do not solve MCQs as intended, as work finds LLMs sans reas…

Question Answering

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

2025-01-06 · CVPR 2025 1 · Yuhui Zhang, Yuchang Su, Yiming Liu, Xiaohan Wang 외

The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluatio…

Language Model EvaluationLanguage ModelingLanguage ModellingMultiple-choice+3

Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

2024-04-16 · Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday 외

Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from…

Ethics

Fantastic Bugs and Where to Find Them in AI Benchmarks

2025-11-20 · Sang Truong, Yuheng Tu, Michael Hardy, Anka Reuel 외 arxiv

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasi…

It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

2025-03-13 · Shrutika Singh, Anton Alyakin, Daniel Alexander Alber, Jaden Stryker 외

The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be…

Multiple-choice