paper-with-me

홈 › Papers

AutoDistill: an End-to-End Framework to Explore and Distill Hardware-Efficient Language Models

2022-01-21 · Xiaofan Zhang, Zongwei Zhou, Deming Chen, Yu Emma Wang

Recently, large pre-trained models have significantly improved the performance of various Natural LanguageProcessing (NLP) tasks but they are expensive to serve due to long serving latency and large memory usage. To compress these models, knowledge distillation has attracted an increasing amount of interest as one of the most effective methods for model compression. However, existing distillation methods have not yet addressed the unique challenges of model serving in datacenters, such as handling fast evolving models, considering serving performance, and optimizing for multiple objectives. To solve these problems, we propose AutoDistill, an end-to-end model distillation framework integrating model architecture exploration and multi-objective optimization for building hardware-efficient NLP pre-trained models. We use Bayesian Optimization to conduct multi-objective Neural Architecture Search for selecting student model architectures. The proposed search comprehensively considers both prediction accuracy and serving latency on target hardware. The experiments on TPUv4i show the finding of seven model architectures with better pre-trained accuracy (up to 3.2% higher) and lower inference latency (up to 1.44x faster) than MobileBERT. By running downstream NLP tasks in the GLUE benchmark, the model distilled for pre-training by AutoDistill with 28.5M parameters achieves an 81.69 average score, which is higher than BERT_BASE, DistillBERT, TinyBERT, NAS-BERT, and MobileBERT. The most compact model found by AutoDistill contains only 20.6M parameters but still outperform BERT_BASE(109M), DistillBERT(67M), TinyBERT(67M), and MobileBERT(25.3M) regarding the average GLUE score. By evaluating on SQuAD, a model found by AutoDistill achieves an 88.4% F1 score with 22.8M parameters, which reduces parameters by more than 62% while maintaining higher accuracy than DistillBERT, TinyBERT, and NAS-BERT.

📄 PDF Abstract BibTeX arXiv:2201.08539

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian OptimizationKnowledge DistillationModel CompressionNeural Architecture Search

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing

2021-09-23 · EMNLP (sustainlp) 2021 11 · Haoyu He, Xingjian Shi, Jonas Mueller, Zha Sheng 외

We aim to identify how different components in the KD pipeline affect the resulting performance and how much the optimal KD pipeline varies across different datasets/tasks, such as the data augmentation policy, the loss …

Data AugmentationHyperparameter Optimization

On Accelerating Edge AI: Optimizing Resource-Constrained Environments

2025-01-25 · Jacob Sander, Achraf Cohen, Venkat R. Dasari, Brent Venable 외

Resource-constrained edge deployments demand AI solutions that balance high performance with stringent compute, memory, and energy limitations. In this survey, we present a comprehensive overview of the primary strategie…

Knowledge DistillationModel CompressionNeural Architecture SearchQuantization+2

Extraction of linearized models from pre-trained networks via knowledge distillation

2026-04-08 · Fumito Kimura, Jun Ohkubo arxiv

Recent developments in hardware, such as photonic integrated circuits and optical devices, are driving demand for research on constructing machine learning architectures tailored for linear operations. Hence, it is valua…

Knowledge Distillation

FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

2026-02-02 · Ruiteng Zhao, Wenshuo Wang, Yicheng Ma, Xiaocong Li 외 arxiv

Force sensing is a crucial modality for Vision-Language-Action (VLA) frameworks, as it enables fine-grained perception and dexterous manipulation in contact-rich tasks. We present Force-Distilled VLA (FD-VLA), a novel fr…

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

2026-07-02 · Junyi Wen, Ruiyan Zhuang, Yongjia Xu, Pengtu Li 외 arxiv

Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints and strict memory hierarchies. While large…

Knowledge Distillation