Papers Multiple-choice
“Multiple-choice” 태그가 달린 논문 1,107편 · 필터 해제
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to answer multiple-choice questions of diff…
Language ModelingLanguage ModellingLarge Language ModelMultiple-choiceHATS: Hindi Analogy Test Set for Evaluating Reasoning in Large Language Models
Analogies test a model's ability to infer implicit relationships between concepts, making them a key benchmark for evaluating reasoning capabilities. While large language models (LLMs) are widely evaluated for reasoning …
Multiple-choiceMateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving,…
FairnessMultiple-choiceAdvanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III
As financial institutions increasingly adopt Large Language Models (LLMs), rigorous domain-specific evaluation becomes critical for responsible deployment. This paper presents a comprehensive benchmark evaluating 23 stat…
Model SelectionMultiple-choiceOmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has severa…
DiversityMultiple-choiceAdapting Vision-Language Models for Evaluating World Models
World models -- generative models that simulate environment dynamics conditioned on past observations and actions -- are gaining prominence in planning, simulation, and embodied AI. However, evaluating their rollouts rem…
Action RecognitionMultimodal ReasoningMultiple-choicePhysUniBench: An Undergraduate-Level Physics Reasoning Benchmark for Multimodal Models
Physics problem-solving is a challenging domain for large AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Current evaluation methodologies sh…
Mathematical ReasoningMultiple-choiceHow Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?
We investigate whether off-the-shelf Multimodal Large Language Models (MLLMs) can tackle Online Episodic-Memory Video Question Answering (OEM-VQA) without additional training. Our pipeline converts a streaming egocentric…
Multiple-choiceQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-lang…
document understandingMultiple-choiceQuestion AnsweringThunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
Negation is a fundamental linguistic phenomenon that poses persistent challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Existing benchmarks often treat negation as …
Multiple-choiceNatural Language InferenceNegationSentenceHypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis …
Decision MakingLanguage ModelingLanguage ModellingMarketing+1Training-free LLM Merging for Multi-task Learning
Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing (NLP) tasks. The release of open-source LLMs like LLaMA and Qwen has triggered the development of numerou…
Multiple-choiceMulti-Task LearningQuestion AnsweringInstruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
Large language models (LLMs) have shown great potential in medical question answering (MedQA), yet adapting them to biomedical reasoning remains challenging due to domain-specific complexity and limited supervision. In t…
Medical Question AnsweringMedQAMultiple-choicePrompt Engineering+1Different Questions, Different Models: Fine-Grained Evaluation of Uncertainty and Calibration in Clinical QA with LLMs
Accurate and well-calibrated uncertainty estimates are essential for deploying large language models (LLMs) in high-stakes domains such as clinical decision support. We present a fine-grained evaluation of uncertainty es…
Multiple-choiceQuestion AnsweringA Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visua…
Multiple-choiceVersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks
Recent advancements in multimodal large language models have successfully extended the Reason-Then-Respond paradigm to image-based reasoning, yet video-based reasoning remains an underdeveloped frontier, primarily due to…
Multiple-choiceOpen-Ended Question AnsweringQuestion AnsweringVideo Captioning+1ARGUS: Hallucination and Omission Evaluation in Video-LLMs
Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questions. Unfortunately, VideoLLMs hallucinat…
DescriptiveFormHallucinationMultiple-choice+2Evaluating LLM-corrupted Crowdsourcing Data Without Ground Truth
The recent success of generative AI highlights the crucial role of high-quality human feedback in building trustworthy AI systems. However, the increasing use of large language models (LLMs) by crowdsourcing workers pose…
Multiple-choiceSTSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
We introduce STSBench, a scenario-based framework to benchmark the holistic understanding of vision-language models (VLMs) for autonomous driving. The framework automatically mines pre-defined traffic scenarios from any …
Autonomous DrivingAutonomous VehiclesDense CaptioningMultiple-choice+2Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
Although vision-language and large language models (VLM and LLM) offer promising opportunities for AI-driven educational assessment, their effectiveness in real-world classroom settings, particularly in underrepresented …
Multiple-choice