Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utterance confidence scores for end-to-end ASR. In those studies, word confidence by itself does not model deletions, and utterance confidence does not take advantage of word-level training signals. This paper proposes to jointly learn word confidence, word deletion, and utterance confidence. Empirical results show that multi-task learning with all three objectives improves confidence metrics (NCE, AUC, RMSE) without the need for increasing the model size of the confidence estimation module. Using the utterance-level confidence for rescoring also decreases the word error rates on Google's Voice Search and Long-tail Maps datasets by 3-5% relative, without needing a dedicated neural rescorer.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task Learningspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Confidence Estimation and Deletion Prediction Using Bidirectional Recurrent Neural Networks
The standard approach to assess reliability of automatic speech transcriptions is through the use of confidence scores. If accurate, these scores provide a flexible mechanism to flag transcription errors for upstream and…
Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text Classification
Recent advances in weakly supervised text classification mostly focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels. In this paper, we revisit the seed matching-based m…
text-classificationText ClassificationAccurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System
Estimating confidence scores for recognition results is a classic task in ASR field and of vital importance for kinds of downstream tasks and training strategies. Previous end-to-end~(E2E) based confidence estimation mod…
speech-recognitionSpeech RecognitionPretraining Chinese BERT for Detecting Word Insertion and Deletion Errors
Chinese BERT models achieve remarkable progress in dealing with grammatical errors of word substitution. However, they fail to handle word insertion and deletion because BERT assumes the existence of a word at each posit…
Language ModelingLanguage ModellingMasked Language ModelingPositionEditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional Fusion
This paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without…
text-to-speechText to Speech