paper-with-me

홈 › Papers

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?

2025-07-21 · Seok Hwan Song, Mohna Chakraborty, Qi Li, Wallapak Tavanapong arxiv

Large Language Models (LLMs) have been evaluated using diverse question types, e.g., multiple-choice, true/false, and short/long answers. This study answers an unexplored question about the impact of different question types on LLM accuracy on reasoning tasks. We investigate the performance of five LLMs on three different types of questions using quantitative and deductive reasoning tasks. The performance metrics include accuracy in the reasoning steps and choosing the final answer. Key Findings: (1) Significant differences exist in LLM performance across different question types. (2) Reasoning accuracy does not necessarily correlate with the final selection accuracy. (3) The number of options and the choice of words, influence LLM performance.

📄 PDF Abstract BibTeX arXiv:2507.15707

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models

2023-12-29 · Yuqing Wang, Yun Zhao

The burgeoning interest in Multimodal Large Language Models (MLLMs), such as OpenAI's GPT-4V(ision), has significantly impacted both academic and industrial realms. These models enhance Large Language Models (LLMs) with …

HellaSwag

Evaluating Mathematical Reasoning Across Large Language Models: A Fine-Grained Approach

2025-03-13 · Afrar Jahin, Arif Hassan Zidan, Wei zhang, Yu Bao 외

With the rapid advancement of Artificial Intelligence (AI), Large Language Models (LLMs) have significantly impacted a wide array of domains, including healthcare, engineering, science, education, and mathematical reason…

Formal LogicMathematical ReasoningMMLU

When Do Program-of-Thoughts Work for Reasoning?

2023-08-29 · Zhen Bi, Ningyu Zhang, Yinuo Jiang, Shumin Deng 외

In the realm of embodied artificial intelligence, the reasoning capabilities of Large Language Models (LLMs) play a pivotal role. Although there are effective methods like program-of-thought prompting for LLMs which uses…

Code GenerationMathematical Reasoning

Zero-Shot Multi-Hop Question Answering via Monte-Carlo Tree Search with Large Language Models

2024-09-28 · Seongmin Lee, Jaewook Shin, Youngjin Ahn, Seokin Seo 외

Recent advances in large language models (LLMs) have significantly impacted the domain of multi-hop question answering (MHQA), where systems are required to aggregate information and infer answers from disparate pieces o…

Multi-hop Question AnsweringQuestion Answering

How Does Quantization Affect Multilingual LLMs?

2024-07-03 · Kelly Marchisio, Saurabh Dash, Hongyu Chen, Dennis Aumiller 외

Quantization techniques are widely used to improve inference speed and deployment of large language models. While a wide body of work examines the impact of quantization on LLMs in English, none have evaluated across lan…

Mathematical ReasoningQuantization