paper-with-me

홈 › Papers

MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models

2023-10-08 · Yifan Wei, Yisong Su, Huanhuan Ma, Xiaoyan Yu, Fangyu Lei, Yuanzhe Zhang, Jun Zhao, Kang Liu

Large language models (LLMs) have shown nearly saturated performance on many natural language processing (NLP) tasks. As a result, it is natural for people to believe that LLMs have also mastered abilities such as time understanding and reasoning. However, research on the temporal sensitivity of LLMs has been insufficiently emphasized. To fill this gap, this paper constructs Multiple Sensitive Factors Time QA (MenatQA), which encompasses three temporal factors (scope factor, order factor, counterfactual factor) with total 2,853 samples for evaluating the time comprehension and reasoning abilities of LLMs. This paper tests current mainstream LLMs with different parameter sizes, ranging from billions to hundreds of billions. The results show most LLMs fall behind smaller temporal reasoning models with different degree on these factors. In specific, LLMs show a significant vulnerability to temporal biases and depend heavily on the temporal information provided in questions. Furthermore, this paper undertakes a preliminary investigation into potential improvement strategies by devising specific prompts and leveraging external tools. These approaches serve as valuable baselines or references for future research endeavors.

📄 PDF Abstract BibTeX arXiv:2310.05157

Code (1)

weiyifan1023/MenatQA 공식 구현 pytorch

Tasks

counterfactual

Similar Papers 제목 키워드 기반

Reasoning with Memory Augmented Neural Networks for Language Comprehension

2016-10-20 · Tsendsuren Munkhdalai, Hong Yu

Hypothesis testing is an important cognitive process that supports human reasoning. In this paper, we introduce a computational hypothesis testing approach based on memory augmented neural networks. Our approach involves…

Reading ComprehensionTwo-sample testing

ESTER: A Machine Reading Comprehension Dataset for Event Semantic Relation Reasoning

2021-04-16 · Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon 외

Understanding how events are semantically related to each other is the essence of reading comprehension. Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations. While th…

Machine Reading ComprehensionNatural Language QueriesQuestion AnsweringReading Comprehension+1

ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic Relations

2021-11-01 · EMNLP 2021 11 · Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon 외

Understanding how events are semantically related to each other is the essence of reading comprehension. Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations. While th…

Machine Reading ComprehensionNatural Language QueriesReading ComprehensionRelation

Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity

2025-11-29 · Subramanyam Sahoo, Vinija Jain, Saanidhya Vats, Siddharth Mohapatra 외 arxiv

Current evaluation of mathematical reasoning in language models relies primarily on answer accuracy, potentially masking fundamental failures in logical computation. We introduce a diagnostic framework that distinguishes…

Mathematical Reasoning

MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models

2024-06-15 · Wentian Wang, Sarthak Jain, Paul Kantor, Jacob Feldman 외

We propose MMLU-SR, a novel dataset designed to measure the true comprehension abilities of Large Language Models (LLMs) by challenging their performance in question-answering tasks with modified terms. We reasoned that …

Mathematical ReasoningMMLUMulti-task Language UnderstandingNatural Language Understanding+1