paper-with-me

홈 › Papers

Distillation Scaling Laws

2025-02-12 · Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb

We provide a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings reduce the risks associated with using distillation at scale; compute allocation for both the teacher and student models can now be done to maximize student performance. We provide compute optimal distillation recipes for when 1) a teacher exists, or 2) a teacher needs training. If many students are to be distilled, or a teacher already exists, distillation outperforms supervised pretraining until a compute level which grows predictably with student size. If one student is to be distilled and a teacher also needs training, supervised learning should be done instead. Additionally, we provide insights across our large scale study of distillation, which increase our understanding of distillation and inform experimental design.

📄 PDF Abstract BibTeX arXiv:2502.08606

Code (0)

등록된 구현이 없습니다.

Tasks

Experimental Design

Similar Papers 제목 키워드 기반

Scaling Laws for Data-Efficient Visual Transfer Learning

2025-04-17 · Wenxuan Yang, Qingqu Wei, Chenxi Ma, Weimin Tan 외

Current scaling laws for visual AI models focus predominantly on large-scale pretraining, leaving a critical gap in understanding how performance scales for data-constrained downstream tasks. To address this limitation, …

Knowledge DistillationTransfer Learning

Scaling Laws for Task-Specific LLM Distillation

2026-06-23 · Lavinia Ghita, Dhruv Desai, Ioana Boier arxiv

Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives e…

General Knowledge

Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment

2025-08-04 · Dai Li, Kevin Course, Wei Li, Hongwei Li 외 arxiv

Scaling laws have been established for recommender systems, yet efficiently deploying foundation model (FM) across multiple recommendation surfaces remains a major unsolved challenge. Existing methods for transfer learni…

Reverse Distillation: Consistently Scaling Protein Language Model Representations

2026-03-08 · Darius Catrina, Christian Bepler, Samuel Sledzieski, Rohit Singh arxiv

Unlike the predictable scaling laws in natural language processing and computer vision, protein language models (PLMs) scale poorly: for many tasks, models within the same family plateau or even decrease in performance, …

Protein Language Model

Unifying Two Types of Scaling Laws from the Perspective of Conditional Kolmogorov Complexity

2025-01-12 · Jun Wan

In 2020, OpenAI proposed the first type of Scaling Laws, describing the relationships between model performance and parameters, data, and compute. In 2024, OpenAI proposed the second type of Scaling Laws, describing the …