paper-with-me

홈 › Papers

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

2024-05-01 · Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, Will Song, Tiffany Zhao, Pranav Raja, Charlotte Zhuang, Dylan Slack, Qin Lyu, Sean Hendryx, Russell Kaplan, Michele Lunati, Summer Yue

Large language models (LLMs) have achieved impressive success on many benchmarks for mathematical reasoning. However, there is growing concern that some of this performance actually reflects dataset contamination, where data closely resembling benchmark questions leaks into the training data, instead of true reasoning ability. To investigate this claim rigorously, we commission Grade School Math 1000 (GSM1k). GSM1k is designed to mirror the style and complexity of the established GSM8k benchmark, the gold standard for measuring elementary mathematical reasoning. We ensure that the two benchmarks are comparable across important metrics such as human solve rates, number of steps in solution, answer magnitude, and more. When evaluating leading open- and closed-source LLMs on GSM1k, we observe accuracy drops of up to 8%, with several families of models showing evidence of systematic overfitting across almost all model sizes. Further analysis suggests a positive relationship (Spearman's r^2 = 0.36) between a model's probability of generating an example from GSM8k and its performance gap between GSM8k and GSM1k, suggesting that some models may have partially memorized GSM8k. Nevertheless, many models, especially those on the frontier, show minimal signs of overfitting, and all models broadly demonstrate generalization to novel math problems guaranteed to not be in their training data.

📄 PDF Abstract BibTeX arXiv:2405.00332

Code (0)

등록된 구현이 없습니다.

Tasks

GSM8KLanguage ModelingLanguage ModellingLarge Language ModelMathMathematical Reasoning

Similar Papers 제목 키워드 기반

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

2023-05-21 · Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying 외

Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be…

SteuerLLM: Local specialized large language model for German tax law analysis

2026-02-11 · Sebastian Wind, Jeta Sopa, Laurin Schmid, Quirin Jackl 외 arxiv

Large language models (LLMs) demonstrate strong general reasoning and language understanding, yet their performance degrades in domains governed by strict formal rules, precise terminology, and legally binding structure.…

Legal Reasoning

A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task

2016-06-09 · ACL 2016 8 · Danqi Chen, Jason Bolton, Christopher D. Manning

Enabling a computer to understand a document so that it can answer comprehension questions is a central, yet unsolved goal of NLP. A key factor impeding its solution by machine learned systems is the limited availability…

ArticlesReading Comprehension

From Phonemes to Meaning: Evaluating Large Language Models on Tamil

2025-11-15 · Jeyarajalingam Varsha, Menan Velayuthan, Sumirtha Karunakaran, Rasan Nivethiga 외 arxiv

Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich languages such as Tamil remains largely…

Self-Verification is All You Need To Pass The Japanese Bar Examination

2026-01-06 · Andrew Shin arxiv

Despite rapid advances in large language models (LLMs), achieving reliable performance on highly professional and structured examinations remains a significant challenge. The Japanese bar examination is a particularly de…

Legal Reasoning