| Rank | Model |
Accuracy | Parameters (Billion) |
Extra Training Data |
Paper | Code | Year |
| 1 |
Claude 3.5 Sonnet (HPT) |
97.72 | – |
|
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
|
devichand579/HPT |
2024 |
| 2 |
DUP prompt upon GPT-4 |
97.1 | – |
|
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
|
whu-zqh/dup |
2024 |
| 3 |
Qwen2-Math-72B-Instruct
(greedy) |
96.7 | 72 |
✓ |
Qwen2 Technical Report
|
qwenlm/qwen1.5 · qwenlm/qwen2 · vicentvankor/sun-shine
· +3 |
2024 |
| 4 |
SFT-Mistral-7B (Metamath, OVM, Smart Ensemble) |
96.4 | 7 |
✓ |
|
|
|
| 5 |
OpenMath2-Llama3.1-70B (majority@256) |
96.0 | – |
✓ |
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
|
NVIDIA/NeMo-Skills |
2024 |
| 6 |
Jiutian-大模型 |
95.2 | 75 |
|
|
|
|
| 7 |
DAMOMath-7B(MetaMath, OVM, BS, Ensemble) |
95.1 | 7 |
✓ |
|
|
|
| 8 |
Claude 3 Opus (0-shot chain-of-thought) |
95 | – |
|
The Claude 3 Model Family: Opus, Sonnet, Haiku
|
|
2024 |
| 9 |
OpenMath2-Llama3.1-70B |
94.9 | – |
✓ |
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
|
NVIDIA/NeMo-Skills |
2024 |
| 10 |
GPT-4 (Teaching-Inspired) |
94.8 | – |
|
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models
|
sallytan13/teaching-inspired-prompting |
2024 |
| 11 |
SFT-Mistral-7B (Metamath + ovm +ensemble) |
94.13 | 7 |
✓ |
|
|
|
| 12 |
OpenMath2-Llama3.1-8B (majority@256) |
94.1 | – |
✓ |
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
|
NVIDIA/NeMo-Skills |
2024 |
| 13 |
Qwen2-72B-Instruct-Step-DPO (0-shot CoT) |
94.0 | – |
✓ |
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
|
dvlab-research/step-dpo |
2024 |
| 14 |
DAMOMath-7B(MetaMath, OVM, Ensemble) |
93.2 | 7 |
✓ |
|
|
|
| 15 |
Claude 3 Sonnet (0-shot chain-of-thought) |
92.3 | – |
|
The Claude 3 Model Family: Opus, Sonnet, Haiku
|
|
2024 |
| 16 |
AlphaLLM (with MCTS) |
92 | 70 |
|
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
|
yetianjhu/alphallm |
2024 |
| 17 |
OpenMath2-Llama3.1-8B |
91.7 | – |
✓ |
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
|
NVIDIA/NeMo-Skills |
2024 |
| 18 |
PaLM 2 (few-shot, k=8, SC) |
91.0 | – |
|
PaLM 2 Technical Report
|
eternityyw/tram-benchmark |
2023 |
| 19 |
GaC(Qwen2-72B-Instruct + Llama-3-70B-Instruct) |
90.91 | – |
|
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
|
yaoching0/gac |
2024 |
| 20 |
OpenMath-CodeLlama-70B (w/ code, SC, k=50) |
90.8 | 70 |
✓ |
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
|
kipok/nemo-skills |
2024 |