paper-with-me

홈 › Papers

EasyMath: A 0-shot Math Benchmark for SLMs

2025-05-20 · Drishya Karki, Michiel Kamphuis, Angelecia Frey

EasyMath is a compact benchmark for practical math reasoning in small language models. It covers thirteen categories, from basic arithmetic and order of operations to word problems, algebraic expressions, edge cases, and omits specialist topics. We tested 23 models (14M to 4B parameters) using exact, numerical, and symbolic checks on free-form answers in a zero-shot setting. Accuracy rises with size and training, chain-of-thought adds modest gains, and consistency improves at scale.

📄 PDF Abstract BibTeX arXiv:2505.14852

Code (0)

등록된 구현이 없습니다.

Tasks

Math

Similar Papers 제목 키워드 기반

AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models

2025-04-30 · Yinghui He, Abhishek Panigrahi, Yong Lin, Sanjeev Arora

In-context learning (ICL) allows a language model to improve its problem-solving capability when provided with suitable information in context. Since the choice of in-context information can be determined based on the pr…

In-Context LearningMath

Beyond Scale: Small Language Models are Comparable to GPT-4 in Mental Health Understanding

2025-07-09 · Hong Jia, Shiya Fu, Feng Xia, Vassilis Kostakos 외 arxiv

The emergence of Small Language Models (SLMs) as privacy-preserving alternatives for sensitive applications raises a fundamental question about their inherent understanding capabilities compared to Large Language Models …

Binary ClassificationFew-Shot Learning

Small or Large? Zero-Shot or Finetuned? Guiding Language Model Choice for Specialized Applications in Healthcare

2025-04-29 · Lovedeep Gondara, Jonathan Simkin, Graham Sayle, Shebnum Devji 외

This study aims to guide language model selection by investigating: 1) the necessity of finetuning versus zero-shot usage, 2) the benefits of domain-adjacent versus generic pretrained models, 3) the value of further doma…

Language ModelingLanguage ModellingModel Selection

HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models

2026-04-14 · Jawad Hossain, Xiangyu Guo, Jiawei Zhou, Chong Liu arxiv

Small language models (SLMs) often struggle with complex mathematical reasoning due to limited capacity to maintain long chains of intermediate steps and to recover from early errors. We address this challenge by introdu…

Mathematical Reasoning

rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

2025-01-08 · Xinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang 외

We present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models. rStar-Math achieves this by exercisi…

Math