paper-with-me

Papers

Benchmarking Large Language Models for Math Reasoning Tasks

2024-08-20 · Kathrin Seßler, Yao Rong, Emek Gözlüklü, Enkelejda Kasneci

The use of Large Language Models (LLMs) in mathematical reasoning has become a cornerstone of related research, demonstrating the intelligence of these models and enabling potential practical applications through their advanced performance, such as in educational settings. Despite the variety of datasets and in-context learning algorithms designed to improve the ability of LLMs to automate mathematical problem solving, the lack of comprehensive benchmarking across different datasets makes it complicated to select an appropriate model for specific tasks. In this project, we present a benchmark that fairly compares seven state-of-the-art in-context learning algorithms for mathematical problem solving across five widely used mathematical datasets on four powerful foundation models. Furthermore, we explore the trade-off between efficiency and performance, highlighting the practical applications of LLMs for mathematical reasoning. Our results indicate that larger foundation models like GPT-4o and LLaMA 3-70B can solve mathematical reasoning independently from the concrete prompting strategy, while for smaller models the in-context learning approach significantly influences the performance. Moreover, the optimal prompt depends on the chosen foundation model. We open-source our benchmark code to support the integration of additional models in future research.

📄 PDF Abstract BibTeX arXiv:2408.10839

Code (1)

kathrinse/math-reasoning-benchmark 공식 구현

Tasks

BenchmarkingIn-Context LearningMathMathematical Problem-SolvingMathematical Reasoning

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

2024-10-06 · Yibo Yan, Shen Wang, Jiahao Huo, Hang Li 외

As the field of Multimodal Large Language Models (MLLMs) continues to evolve, their potential to revolutionize artificial intelligence is particularly promising, especially in addressing mathematical reasoning tasks. Cur…

BenchmarkingMathematical Reasoning

ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian

2025-01-12 · Mykyta Syromiatnikov, Victoria Ruvinskaya, Anastasiya Troynina

As the usage of large language models for problems outside of simple text understanding or generation increases, assessing their abilities and limitations becomes crucial. While significant progress has been made in this…

BenchmarkingMathMultiple-choice

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

2025-05-21 · Chengwei Wei, Bin Wang, Jung-jae Kim, Nancy F. Chen

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have led to strong reasoning ability across a wide range of tasks. However, their ability to perform mathematical reasoning from spoken input re…

BenchmarkingMathMathematical Problem-SolvingMathematical Reasoning+1

MathRobust-LV: Evaluation of Large Language Models' Robustness to Linguistic Variations in Mathematical Reasoning

2025-10-07 · Neeraja Kirtane, Yuvraj Khanna, Peter Relan arxiv

Large language models excel on math benchmarks, but their math reasoning robustness to linguistic variation is underexplored. While recent work increasingly treats high-difficulty competitions like the IMO as the gold st…

Mathematical Reasoning

MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data

2024-06-26 · Meng Fang, Xiangpeng Wan, Fei Lu, Fei Xing 외

Large language models (LLMs) have significantly advanced natural language understanding and demonstrated strong problem-solving abilities. Despite these successes, most LLMs still struggle with solving mathematical probl…

BenchmarkingMathMathematical Problem-SolvingMathematical Reasoning+1