On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
Internal language model (ILM) subtraction has been widely applied to improve the performance of the RNN-Transducer with external language model (LM) fusion for speech recognition. In this work, we show that sequence discriminative training has a strong correlation with ILM subtraction from both theoretical and empirical points of view. Theoretically, we derive that the global optimum of maximum mutual information (MMI) training shares a similar formula as ILM subtraction. Empirically, we show that ILM subtraction and sequence discriminative training achieve similar effects across a wide range of experiments on Librispeech, including both MMI and minimum Bayes risk (MBR) criteria, as well as neural transducers and LMs of both full and limited context. The benefit of ILM subtraction also becomes much smaller after sequence discriminative training. We also provide an in-depth study to show that sequence discriminative training has a minimal effect on the commonly used zero-encoder ILM estimation, but a joint effect on both encoder and prediction + joint network for posterior probability reshaping including both ILM and blank suppression.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingRelationspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Graph-propagation based Correlation Learning for Weakly Supervised Fine-grained Image Classification
The key of Weakly Supervised Fine-grained Image Classification (WFGIC) is how to pick out the discriminative regions and learn the discriminative features from them. However, most recent WFGIC methods pick out the discri…
Fine-Grained Image ClassificationGeneral Classificationimage-classificationImage ClassificationLearning Discriminative Relational Features for Sequence Labeling
Discovering relational structure between input features in sequence labeling models has shown to improve their accuracy in several problem settings. However, the search space of relational features is exponential in the …
Self-Contextualized Attention for Abusive Language Identification
The use of attention mechanisms in deep learning approaches has become popular in natural language processing due to its outstanding performance. The use of these mechanisms allows one managing the importance of the elem…
Abusive LanguageLanguage IdentificationMCNS: Mining Causal Natural Structures Inside Time Series via A Novel Internal Causality Scheme
Causal inference permits us to discover covert relationships of various variables in time series. However, in most existing works, the variables mentioned above are the dimensions. The causality between dimensions could …
Causal InferenceTime SeriesTime Series ClassificationLow-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
We propose a neural language modeling system based on low-rank adaptation (LoRA) for speech recognition output rescoring. Although pretrained language models (LMs) like BERT have shown superior performance in second-pass…
Language ModelingLanguage ModellingLarge Language Modelspeech-recognition+1