paper-with-me

홈 › Papers

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

2025-10-06 · Sara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero, Francesco Locatello, David Alvarez-Melis arxiv

Large language models (LLMs) are typically deployed under diverse memory and compute constraints. Existing approaches build model families by training each size independently, which is prohibitively expensive and provides only coarse-grained size options. In this work, we identify a novel phenomenon that we call boomerang distillation: starting from a large base model (the teacher), one first distills down to a small student and then progressively reconstructs intermediate-sized models by re-incorporating blocks of teacher layers into the student without any additional training. This process produces zero-shot interpolated models of many intermediate sizes whose performance scales smoothly between the student and teacher, often matching or surpassing pretrained or distilled models of the same size. We further analyze when this type of interpolation succeeds, showing that alignment between teacher and student through pruning and distillation is essential. Boomerang distillation thus provides a simple and efficient way to generate fine-grained model families, dramatically reducing training cost while enabling flexible adaptation across deployment environments. The code and models are available at https://github.com/dcml-lab/boomerang-distillation.

📄 PDF Abstract BibTeX arXiv:2510.05064

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Evolution of Boomerang Uniformity in Cryptographic S-boxes

2022-12-09 · Marko Djurasevic, Domagoj Jakobovic, Luca Mariot, Sihem Mesnager 외

S-boxes are an important primitive that help cryptographic algorithms to be resilient against various attacks. The resilience against specific attacks can be connected with a certain property of an S-box, and the better …

Understanding Layer Patching in Model Size Interpolation

2026-07-09 · Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak, Marco Fumero 외 arxiv

Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training. Recent work on boomerang distillation [Kangaslahti et al., 2026] shows t…

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

2026-08-24 · Yan Zhou, Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak 외 arxiv

Practical deployment of large language models (LLMs) requires families of post-trained variants---instruction-tuned, reasoning-tuned, and chat-style models---each at multiple sizes to meet diverse latency and memory budg…

Boomerang: Local sampling on image manifolds using diffusion models

2022-10-21 · Lorenzo Luzi, Paul M Mayer, Josue Casco-Rodriguez, Ali Siahkoohi 외

The inference stage of diffusion models can be seen as running a reverse-time diffusion stochastic differential equation, where samples from a Gaussian latent distribution are transformed into samples from a target distr…

Data AugmentationImage EnhancementImage Super-ResolutionPrivacy Preserving+2

Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition

2025-05-23 · Minxue Niu, Emily Mower Provost

The ability to handle various emotion labels without dedicated training is crucial for building adaptable Emotion Recognition (ER) systems. Conventional ER models rely on training using fixed label sets and struggle to g…

DescriptiveEmotion Recognition