paper-with-me

Papers Multiple-choice

“Multiple-choice” 태그가 달린 논문 1,107편 · 필터 해제

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations

2025-07-17 · Carlos Arriaga, Gonzalo Martínez, Eneko Sendin, Javier Conde 외

The evaluation of large language models is a complex task, in which several approaches have been proposed. The most common is the use of automated benchmarks in which LLMs have to answer multiple-choice questions of diff…

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice

HATS: Hindi Analogy Test Set for Evaluating Reasoning in Large Language Models

2025-07-17 · Ashray Gupta, Rohan Joseph, Sunny Rai

Analogies test a model's ability to infer implicit relationships between concepts, making them a key benchmark for evaluating reasoning capabilities. While large language models (LLMs) are widely evaluated for reasoning …

Multiple-choice

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

2025-07-03 · Dumitran Adrian Marius, Theodor-Pierre Moroianu, Buca Mihnea-Vicentiu

The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving,…

FairnessMultiple-choice

Advanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III

2025-06-29 · Pranam Shetty, Abhisek Upadhayaya, Parth Mitesh Shah, Srikanth Jagabathula 외

As financial institutions increasingly adopt Large Language Models (LLMs), rigorous domain-specific evaluation becomes critical for responsible deployment. This paper presents a comprehensive benchmark evaluating 23 stat…

Model SelectionMultiple-choice

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

2025-06-26 · Yiman Zhang, Ziheng Luo, Qiangyu Yan, wei he 외

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has severa…

DiversityMultiple-choice

Adapting Vision-Language Models for Evaluating World Models

2025-06-22 · Mariya Hendriksen, Tabish Rashid, David Bignell, Raluca Georgescu 외

World models -- generative models that simulate environment dynamics conditioned on past observations and actions -- are gaining prominence in planning, simulation, and embodied AI. However, evaluating their rollouts rem…

Action RecognitionMultimodal ReasoningMultiple-choice

PhysUniBench: An Undergraduate-Level Physics Reasoning Benchmark for Multimodal Models

2025-06-21 · Lintao Wang, Encheng Su, Jiaqi Liu, Pengze Li 외

Physics problem-solving is a challenging domain for large AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Current evaluation methodologies sh…

Mathematical ReasoningMultiple-choice

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?

2025-06-19 · Giuseppe Lando, Rosario Forte, Giovanni Maria Farinella, Antonino Furnari

We investigate whether off-the-shelf Multimodal Large Language Models (MLLMs) can tackle Online Episodic-Memory Video Question Answering (OEM-VQA) without additional training. Our pipeline converts a streaming egocentric…

Multiple-choiceQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)

WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts

2025-06-18 · Negar Foroutan, Angelika Romanou, Matin Ansaripour, Julian Martin Eisenschlos 외

Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-lang…

document understandingMultiple-choiceQuestion Answering

Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding

2025-06-17 · Yeonkyoung So, Gyuseong Lee, Sungmok Jung, Joonhak Lee 외

Negation is a fundamental linguistic phenomenon that poses persistent challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Existing benchmarks often treat negation as …

Multiple-choiceNatural Language InferenceNegationSentence

Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings

2025-06-17 · Harbin Hong, Sebastian Caldas, Liu Leqi

As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis …

Decision MakingLanguage ModelingLanguage ModellingMarketing+1

Training-free LLM Merging for Multi-task Learning

2025-06-14 · Zichuan Fu, Xian Wu, Yejing Wang, Wanyu Wang 외

Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing (NLP) tasks. The release of open-source LLMs like LLaMA and Qwen has triggered the development of numerou…

Multiple-choiceMulti-Task LearningQuestion Answering

Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs

2025-06-13 · Chenqian Le, Ziheng Gong, Chihang Wang, Haowei Ni 외

Large language models (LLMs) have shown great potential in medical question answering (MedQA), yet adapting them to biomedical reasoning remains challenging due to domain-specific complexity and limited supervision. In t…

Medical Question AnsweringMedQAMultiple-choicePrompt Engineering+1

Different Questions, Different Models: Fine-Grained Evaluation of Uncertainty and Calibration in Clinical QA with LLMs

2025-06-12 · Alberto Testoni, Iacer Calixto

Accurate and well-calibrated uncertainty estimates are essential for deploying large language models (LLMs) in high-stakes domains such as clinical decision support. We present a fine-grained evaluation of uncertainty es…

Multiple-choiceQuestion Answering

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

2025-06-11 · Benno Krojer, Mojtaba Komeili, Candace Ross, Quentin Garrido 외

Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visua…

Multiple-choice

VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks

2025-06-10 · Xinlong Chen, Yuanxing Zhang, Yushuo Guan, Bohan Zeng 외

Recent advancements in multimodal large language models have successfully extended the Reason-Then-Respond paradigm to image-based reasoning, yet video-based reasoning remains an underdeveloped frontier, primarily due to…

Multiple-choiceOpen-Ended Question AnsweringQuestion AnsweringVideo Captioning+1

ARGUS: Hallucination and Omission Evaluation in Video-LLMs

2025-06-09 · Ruchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli 외

Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questions. Unfortunately, VideoLLMs hallucinat…

DescriptiveFormHallucinationMultiple-choice+2

Evaluating LLM-corrupted Crowdsourcing Data Without Ground Truth

2025-06-08 · Yichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang Liu

The recent success of generative AI highlights the crucial role of high-quality human feedback in building trustworthy AI systems. However, the increasing use of large language models (LLMs) by crowdsourcing workers pose…

Multiple-choice

STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

2025-06-06 · Christian Fruhwirth-Reisinger, Dušan Malić, Wei Lin, David Schinagl 외

We introduce STSBench, a scenario-based framework to benchmark the holistic understanding of vision-language models (VLMs) for autonomous driving. The framework automatically mines pre-defined traffic scenarios from any …

Autonomous DrivingAutonomous VehiclesDense CaptioningMultiple-choice+2

Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms

2025-06-05 · Nurul Aisyah, Muhammad Dehan Al Kautsar, Arif Hidayat, Raqib Chowdhury 외

Although vision-language and large language models (VLM and LLM) offer promising opportunities for AI-driven educational assessment, their effectiveness in real-world classroom settings, particularly in underrepresented …

Multiple-choice
1–20 / 1,107 다음 →