paper-with-me

홈 › Papers

Knowledge Distillation for Small-footprint Highway Networks

2016-08-02 · Liang Lu, Michelle Guo, Steve Renals

Deep learning has significantly advanced state-of-the-art of speech recognition in the past few years. However, compared to conventional Gaussian mixture acoustic models, neural network models are usually much larger, and are therefore not very deployable in embedded devices. Previously, we investigated a compact highway deep neural network (HDNN) for acoustic modelling, which is a type of depth-gated feedforward neural network. We have shown that HDNN-based acoustic models can achieve comparable recognition accuracy with much smaller number of model parameters compared to plain deep neural network (DNN) acoustic models. In this paper, we push the boundary further by leveraging on the knowledge distillation technique that is also known as {\it teacher-student} training, i.e., we train the compact HDNN model with the supervision of a high accuracy cumbersome model. Furthermore, we also investigate sequence training and adaptation in the context of teacher-student training. Our experiments were performed on the AMI meeting speech recognition corpus. With this technique, we significantly improved the recognition accuracy of the HDNN acoustic model with less than 0.8 million parameters, and narrowed the gap between this model and the plain DNN with 30 million parameters.

📄 PDF Abstract BibTeX arXiv:1608.00892

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic ModellingKnowledge Distillationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Small-footprint Deep Neural Networks with Highway Connections for Speech Recognition

2015-12-14 · Liang Lu, Steve Renals

For speech recognition, deep neural networks (DNNs) have significantly improved the recognition accuracy in most of benchmark datasets and application domains. However, compared to the conventional Gaussian mixture model…

speech-recognitionSpeech Recognition

Life Cycle-Aware Evaluation of Knowledge Distillation for Machine Translation: Environmental Impact and Translation Quality Trade-offs

2026-02-10 · Joseph Attieh, Timothee Mickus, Anne-Laure Ligozat, Aurélie Névéol 외 arxiv

Knowledge distillation (KD) is a tool to compress a larger system (teacher) into a smaller one (student). In machine translation, studies typically report only the translation quality of the student and omit the computat…

Knowledge DistillationMachine Translation

Extremely Small BERT Models from Mixed-Vocabulary Training

2019-09-25 · EACL 2021 2 · Sanqiang Zhao, Raghav Gupta, Yang song, Denny Zhou

Pretrained language models like BERT have achieved good results on NLP tasks, but are impractical on resource-limited devices due to memory footprint. A large fraction of this footprint comes from the input embeddings wi…

Knowledge DistillationLanguage ModellingModel CompressionWord Embeddings

Small-footprint Highway Deep Neural Networks for Speech Recognition

2016-10-18 · Liang Lu, Steve Renals

State-of-the-art speech recognition systems typically employ neural network acoustic models. However, compared to Gaussian mixture models, deep neural network (DNN) based acoustic models often have many more model parame…

speech-recognitionSpeech Recognition

M2KD: Multi-model and Multi-level Knowledge Distillation for Incremental Learning

2019-04-03 · Peng Zhou, Long Mai, Jianming Zhang, Ning Xu 외

Incremental learning targets at achieving good performance on new categories without forgetting old ones. Knowledge distillation has been shown critical in preserving the performance on old classes. Conventional methods,…

Incremental LearningKnowledge Distillation