paper-with-me

홈 › Papers

Utterance-level neural confidence measure for end-to-end children speech recognition

2021-09-16 · Wei Liu, Tan Lee

Confidence measure is a performance index of particular importance for automatic speech recognition (ASR) systems deployed in real-world scenarios. In the present study, utterance-level neural confidence measure (NCM) in end-to-end automatic speech recognition (E2E ASR) is investigated. The E2E system adopts the joint CTC-attention Transformer architecture. The prediction of NCM is formulated as a task of binary classification, i.e., accept/reject the input utterance, based on a set of predictor features acquired during the ASR decoding process. The investigation is focused on evaluating and comparing the efficacies of predictor features that are derived from different internal and external modules of the E2E system. Experiments are carried out on children speech, for which state-of-the-art ASR systems show less than satisfactory performance and robust confidence measure is particularly useful. It is noted that predictor features related to acoustic information of speech play a more important role in estimating confidence measure than those related to linguistic information. N-best score features show significantly better performance than single-best ones. It has also been shown that the metrics of EER and AUC are not appropriate to evaluate the NCM of a mismatched ASR with significant performance gap.

📄 PDF Abstract BibTeX arXiv:2109.07750

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Binary Classificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
Residual Connection 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Speech Timing in Typically Developing Mandarin-Speaking Children From Ages 3 To 4

2022-11-01 · ROCLING 2022 11 · Jeng Man Lew, Li-mei Chen, Yu Ching Lin

This study aims to develop a better understanding of the speech timing development in Mandarin-speaking children from 3 to 4 years of age. Data were selected from two typically developing children. Four 50-min recordings…

Transcribing Children's Speech: ASR Performance and Obtaining Reliable Orthographic Transcriptions

2026-04-10 · Gus Lathouwers, Lingyun Gao, Catia Cucchiarini, Helmer Strik arxiv

Automatic speech recognition (ASR) has the potential to substantially reduce manual annotation effort in child speech research by generating automatic transcriptions. However, obtaining reliably high-quality ASR transcri…

Speech Recognition

speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment

2021-04-03 · Junbo Zhang, Zhiwen Zhang, Yongqing Wang, Zhiyong Yan 외

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are c…

Phone-level pronunciation scoringSentencespeech-recognition

Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction

2021-04-26 · David Qiu, Yanzhang He, Qiujia Li, Yu Zhang 외

Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utterance confidence scores for end-to-end AS…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task Learningspeech-recognition+1

End-to-End Neural Systems for Automatic Children Speech Recognition: An Empirical Study

2021-02-19 · Prashanth Gurunath Shivakumar, Shrikanth Narayanan

A key desiderata for inclusive and accessible speech recognition technology is ensuring its robust performance to children's speech. Notably, this includes the rapidly advancing neural network based end-to-end speech rec…

speech-recognitionSpeech Recognition