paper-with-me

홈 › Papers

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

2023-09-11 · Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, Wenhu Chen

We introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, our meticulously curated instruction tuning dataset. MathInstruct is compiled from 13 math datasets with intermediate rationales, six of which have rationales newly curated by us. It presents a unique hybrid of chain-of-thought (CoT) and program-of-thought (PoT) rationales, and also ensures extensive coverage of diverse fields in math. The hybrid of CoT and PoT not only unleashes the potential of tool use but also allows different thought processes for different math problems. As a result, the MAmmoTH series substantially outperform existing open-source models on nine mathematical reasoning datasets across all scales with an average accuracy gain between 16% and 32%. Remarkably, our MAmmoTH-7B model reaches 33% on MATH (a competition-level dataset), which exceeds the best open-source 7B model (WizardMath) by 23%, and the MAmmoTH-34B model achieves 44% accuracy on MATH, even surpassing GPT-4's CoT result. Our work underscores the importance of diverse problem coverage and the use of hybrid rationales in developing superior math generalist models.

📄 PDF Abstract BibTeX arXiv:2309.05653

Code (1)

tiger-ai-lab/mammoth

Tasks

MathMathematical Reasoning

Similar Papers 제목 키워드 기반

MAmmoTH2: Scaling Instructions from the Web

2024-05-06 · Xiang Yue, Tuney Zheng, Ge Zhang, Wenhu Chen

Instruction tuning improves the reasoning abilities of large language models (LLMs), with data quality and scalability being the crucial factors. Most instruction tuning data come from human crowd-sourcing or GPT-4 disti…

ChatbotGSM8KMath

Assessing the Emergent Symbolic Reasoning Abilities of Llama Large Language Models

2024-06-05 · Flavio Petruzzellis, Alberto Testolin, Alessandro Sperduti

Large Language Models (LLMs) achieve impressive performance in a wide range of tasks, even if they are often trained with the only objective of chatting fluently with users. Among other skills, LLMs show emergent abiliti…

Mathematical Reasoning

Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks

2024-04-19 · Avinash Anand, Mohit Gupta, Kritarth Prasad, Navya Singla 외

The rapid progress in the field of natural language processing (NLP) systems and the expansion of large language models (LLMs) have opened up numerous opportunities in the field of education and instructional methods. Th…

Mathematical Problem-Solving

Excess of genomic defects in a woolly mammoth on Wrangel island

2017-01-19

Woolly mammoths (Mammuthus primigenius) populated Siberia, Beringia, and North America during the Pleistocene and early Holocene. Recent breakthroughs in ancient DNA sequencing have allowed for complete genome sequencing…

VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search

2025-03-13 · Yiming Jia, Jiachen Li, Xiang Yue, Bo Li 외

Vision-Language Models have made significant progress on many perception-focused tasks, however, their progress on reasoning-focused tasks seem to be limited due to the lack of high-quality and diverse training data. In …

Image RetrievalMath