paper-with-me

Papers

FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models

2024-03-12 · Yan Liu, Renren Jin, Ling Shi, Zheng Yao, Deyi Xiong

To thoroughly assess the mathematical reasoning abilities of Large Language Models (LLMs), we need to carefully curate evaluation datasets covering diverse mathematical concepts and mathematical problems at different difficulty levels. In pursuit of this objective, we propose FineMath in this paper, a fine-grained mathematical evaluation benchmark dataset for assessing Chinese LLMs. FineMath is created to cover the major key mathematical concepts taught in elementary school math, which are further divided into 17 categories of math word problems, enabling in-depth analysis of mathematical reasoning abilities of LLMs. All the 17 categories of math word problems are manually annotated with their difficulty levels according to the number of reasoning steps required to solve these problems. We conduct extensive experiments on a wide range of LLMs on FineMath and find that there is still considerable room for improvements in terms of mathematical reasoning capability of Chinese LLMs. We also carry out an in-depth analysis on the evaluation process and methods that have been overlooked previously. These two factors significantly influence the model results and our understanding of their mathematical reasoning capabilities. The dataset will be publicly available soon.

📄 PDF Abstract BibTeX arXiv:2403.07747

Code (0)

등록된 구현이 없습니다.

Tasks

MathMathematical Reasoning

Similar Papers 제목 키워드 기반

Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset

2025-08-20 · Rabeeh Karimi Mahabadi, Sanjeev Satheesh, Shrimai Prabhumoye, Mostofa Patwary 외 arxiv

Pretraining large language models (LLMs) on high-quality, structured data such as mathematics and code substantially enhances reasoning capabilities. However, existing math-focused datasets built from Common Crawl suffer…

MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning

2025-07-24 · Xiaoyuan Li, Moxin Li, Wenjie Wang, Rui Men 외 arxiv

Recent progress in Multi-modal Large Language Models (MLLMs) has enabled step-by-step multi-modal mathematical reasoning by performing visual operations based on the textual instructions. A promising approach uses code a…

Mathematical ReasoningCode Generation

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

2024-06-28 · Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin, Zhi-Long Ji 외

Due to the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Despite the datasets like MathVista proposed benchmarks for as…

DiversityMath

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

2026-06-29 · Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov 외 arxiv

As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right re…

Information Retrieval

ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models

2024-02-22 · Yanan Wu, Jie Liu, Xingyuan Bu, Jiaheng Liu 외

This paper introduces ConceptMath, a bilingual (English and Chinese), fine-grained benchmark that evaluates concept-wise mathematical reasoning of Large Language Models (LLMs). Unlike traditional benchmarks that evaluate…

MathMathematical Reasoning