paper-with-me

Papers

Small Language Models (SLMs) Can Still Pack a Punch: A survey

2025-01-03 · Shreyas Subramanian, Vikram Elango, Mecit Gungor

As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 160 papers presents a family of Small Language Models (SLMs) in the 1 to 8 billion parameter range that demonstrate smaller models can perform as well, or even outperform large models. We explore task agnostic, general purpose SLMs, task-specific SLMs and techniques to create SLMs that can guide the community to build models while balancing performance, efficiency, scalability and cost. Furthermore we define and characterize SLMs' effective sizes, representing increased capability with respect to LLMs.

📄 PDF Abstract BibTeX arXiv:2501.05465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models

2023-11-15 · Weize Liu, Guocong Li, Kai Zhang, Bang Du 외

Large language models (LLMs) have achieved remarkable advancements in natural language processing. However, the massive scale and computational demands of these models present formidable challenges when considering their…

Transfer Learning

Distilling Empathy from Large Language Models

2025-07-10 · Henry J. Xie, Jinghan Zhang, Xinhao Zhang, Kunpeng Liu arxiv

The distillation of knowledge from Large Language Models (LLMs) into Smaller Language Models (SLMs), preserving the capabilities and performance of LLMs while reducing model size, has played a key role in the proliferati…

Distilling Mathematical Reasoning Capabilities into Small Language Models

2024-01-22 · Xunyu Zhu, Jian Li, Yong liu, Can Ma 외

This work addresses the challenge of democratizing advanced Large Language Models (LLMs) by compressing their mathematical reasoning capabilities into sub-billion parameter Small Language Models (SLMs) without compromisi…

Mathematical Reasoning

Distilling LLM Agent into Small Models with Retrieval and Code Tools

2025-05-23 · Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho 외

Large language models (LLMs) excel at complex reasoning tasks but remain computationally expensive, limiting their practical deployment. To address this, recent works have focused on distilling reasoning capabilities int…

Action GenerationDomain GeneralizationRetrieval

Distilling Fine-grained Sentiment Understanding from Large Language Models

2024-12-24 · Yice Zhang, Guangyu Xie, Hongling Xu, Kaiheng Hou 외

Fine-grained sentiment analysis (FSA) aims to extract and summarize user opinions from vast opinionated text. Recent studies demonstrate that large language models (LLMs) possess exceptional sentiment understanding capab…

Sentiment AnalysisSentiment ClassificationZero-shot Sentiment Classification