paper-with-me

Papers

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

2024-05-20 · Hongwei Liu, Zilong Zheng, Yuxuan Qiao, Haodong Duan, Zhiwei Fei, Fengzhe Zhou, Wenwei Zhang, Songyang Zhang, Dahua Lin, Kai Chen

Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional perspective, falling short in providing a holistic assessment of the LLMs' math capabilities. To address this gap, we introduce MathBench, a new benchmark that rigorously assesses the mathematical capabilities of large language models. MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills. The benchmark progresses through five distinct stages, from basic arithmetic to college mathematics, and is structured to evaluate models at various depths of knowledge. Each stage includes theoretical questions and application problems, allowing us to measure a model's mathematical proficiency and its ability to apply concepts in practical scenarios. MathBench aims to enhance the evaluation of LLMs' mathematical abilities, providing a nuanced view of their knowledge understanding levels and problem solving skills in a bilingual context. The project is released at https://github.com/open-compass/MathBench .

📄 PDF Abstract BibTeX arXiv:2405.12209

Code (1)

open-compass/mathbench 공식 구현 pytorch

Tasks

College MathematicsGSM8KMath

Similar Papers 제목 키워드 기반

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

2025-01-23 · Xin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao 외

Large Language Models (LLMs) have made significant strides in mathematical reasoning, underscoring the need for a comprehensive and fair evaluation of their capabilities. However, existing benchmarks often fall short, ei…

Mathematical Reasoning

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

2026-06-02 · Zetian Ouyang, Linlin Wang, Gerard de Melo, Liang He arxiv

Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and ma…

Mathematical Reasoning

BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios

2026-02-19 · Yunseung Lee, Subin Kim, Youngjun Kwak, Jaegul Choo arxiv

Large language models (LLMs)-based chatbots are increasingly being adopted in the financial domain, particularly in digital banking, to handle customer inquiries about products such as deposits, savings, and loans. Howev…

WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications

2025-05-20 · Xin Li, Mengbing Liu, Li Wei, Jiancheng An 외

Large Language Models (LLMs) have achieved impressive results across a broad array of tasks, yet their capacity for complex, domain-specific mathematical reasoning-particularly in wireless communications-remains underexp…

Mathematical ReasoningMultiple-choice

StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error

2025-03-13 · Shu-Xun Yang, Cunxiang Wang, Yidong Wang, Xiaotao Gu 외

Evaluating mathematical capabilities is critical for assessing the overall performance of large language models (LLMs). However, existing evaluation methods often focus solely on final answers, resulting in highly inaccu…

Math