paper-with-me

홈 › Papers

Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech Recognition

2019-07-13 · Ye Bai, Jiangyan Yi, Jian-Hua Tao, Zhengkun Tian, Zhengqi Wen

Integrating an external language model into a sequence-to-sequence speech recognition system is non-trivial. Previous works utilize linear interpolation or a fusion network to integrate external language models. However, these approaches introduce external components, and increase decoding computation. In this paper, we instead propose a knowledge distillation based training approach to integrating external language models into a sequence-to-sequence model. A recurrent neural network language model, which is trained on large scale external text, generates soft labels to guide the sequence-to-sequence model training. Thus, the language model plays the role of the teacher. This approach does not add any external component to the sequence-to-sequence model during testing. And this approach is flexible to be combined with shallow fusion technique together for decoding. The experiments are conducted on public Chinese datasets AISHELL-1 and CLMAD. Our approach achieves a character error rate of 9.3%, which is relatively reduced by 18.42% compared with the vanilla sequence-to-sequence model.

📄 PDF Abstract BibTeX arXiv:1907.06017

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationLanguage ModelingLanguage ModellingSequence-To-Sequence Speech Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Born Again Neural Networks

2018-05-12 · ICML 2018 7 · Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti 외

Knowledge distillation (KD) consists of transferring knowledge from one machine learning model (the teacher}) to another (the student). Commonly, the teacher is a high-capacity model with formidable performance, while th…

Image ClassificationKnowledge DistillationLanguage Modeling

Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?

2022-12-16 · Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang 외

The success of deep learning heavily relies on large-scale data with comprehensive labels, which is more expensive and time-consuming to fetch in 3D compared to 2D images or natural languages. This promotes the potential…

3D Point Cloud ClassificationFew-Shot 3D Point Cloud ClassificationKnowledge DistillationRepresentation Learning

Distillation from Heterogeneous Models for Top-K Recommendation

2023-03-02 · SeongKu Kang, Wonbin Kweon, Dongha Lee, Jianxun Lian 외

Recent recommender systems have shown remarkable performance by using an ensemble of heterogeneous models. However, it is exceedingly costly because it requires resources and inference latency proportional to the number …

Knowledge DistillationRecommendation SystemsTransfer Learning

AuG-KD: Anchor-Based Mixup Generation for Out-of-Domain Knowledge Distillation

2024-03-11 · Zihao Tang, Zheqi Lv, Shengyu Zhang, Yifan Zhou 외

Due to privacy or patent concerns, a growing number of large models are released without granting access to their training data, making transferring their knowledge inefficient and problematic. In response, Data-Free Kno…

Data-free Knowledge DistillationKnowledge Distillation

Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data

2019-12-04 · Ye Bai, Jiangyan Yi, Jian-Hua Tao, Zhengqi Wen 외

Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained with speech-text paired data. It is cha…

Language ModellingSentenceSequence-To-Sequence Speech Recognitionspeech-recognition+1