paper-with-me

홈 › Papers

Common 7B Language Models Already Possess Strong Math Capabilities

2024-03-07 · Chen Li, Weiqi Wang, Jingcheng Hu, Yixuan Wei, Nanning Zheng, Han Hu, Zheng Zhang, Houwen Peng

Mathematical capabilities were previously believed to emerge in common language models only at a very large scale or require extensive math-related pre-training. This paper shows that the LLaMA-2 7B model with common pre-training already exhibits strong mathematical abilities, as evidenced by its impressive accuracy of 97.7% and 72.0% on the GSM8K and MATH benchmarks, respectively, when selecting the best response from 256 random generations. The primary issue with the current base model is the difficulty in consistently eliciting its inherent mathematical capabilities. Notably, the accuracy for the first answer drops to 49.5% and 7.9% on the GSM8K and MATH benchmarks, respectively. We find that simply scaling up the SFT data can significantly enhance the reliability of generating correct answers. However, the potential for extensive scaling is constrained by the scarcity of publicly available math questions. To overcome this limitation, we employ synthetic data, which proves to be nearly as effective as real data and shows no clear saturation when scaled up to approximately one million samples. This straightforward approach achieves an accuracy of 82.6% on GSM8K and 40.6% on MATH using LLaMA-2 7B models, surpassing previous models by 14.2% and 20.8%, respectively. We also provide insights into scaling behaviors across different reasoning complexities and error types.

📄 PDF Abstract BibTeX arXiv:2403.04706

Code (2)

xwin-lm/xwin-lm 공식 구현 pytorch
jerrywu-code/susgen pytorch

Tasks

GSM8KMath

Methods 이 논문이 사용한 방법론

SFT Shrink and Fine-Tune, or SFT, is a type of distillation that avoids explicit distillation by copying parameters to a student student model and then fine-tuning.…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Measuring and Improving BERT's Mathematical Abilities by Predicting the Order of Reasoning.

2021-08-01 · ACL 2021 5 · Piotr Pi{\k{e}}kos, Mateusz Malinowski, Henryk Michalewski

Imagine you are in a supermarket. You have two bananas in your basket and want to buy four apples. How many fruits do you have in total? This seemingly straightforward question can be challenging for data-driven language…

Language ModelingLanguage ModellingMath

Measuring and Improving BERT's Mathematical Abilities by Predicting the Order of Reasoning

2021-06-07 · Piotr Piękos, Henryk Michalewski, Mateusz Malinowski

Imagine you are in a supermarket. You have two bananas in your basket and want to buy four apples. How many fruits do you have in total? This seemingly straightforward question can be challenging for data-driven language…

Language ModelingLanguage ModellingMath

Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers

2025-05-26 · Rihui Xin, Han Liu, Zecheng Wang, Yupeng Zhang 외

Large Language Models have achieved remarkable success in natural language processing tasks, with Reinforcement Learning playing a key role in adapting them to specific applications. However, obtaining ground truth answe…

Logical ReasoningMathematical Problem-Solving

Graph Convolutional Neural Networks as Parametric CoKleisli morphisms

2022-12-01 · Bruno Gavranović, Mattia Villani

We define the bicategory of Graph Convolutional Neural Networks $\mathbf{GCNN}_n$ for an arbitrary graph with $n$ nodes. We show it can be factored through the already existing categorical constructions for deep learning…

Inductive Bias

Unexplored flaws in multiple-choice VQA evaluations

2025-11-27 · Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagodonat 외 arxiv

Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in handling image-text inputs. A common way to assess this ability is through multiple-choice Visual Question Answering (VQA). Earlier works have a…

Visual Question Answering