Weakly-supervised text-to-speech alignment confidence measure
This work proposes a new confidence measure for evaluating text-to-speech alignment systems outputs, which is a key component for many applications, such as semi-automatic corpus anonymization, lips syncing, film dubbing, corpus preparation for speech synthesis and speech recognition acoustic models training. This confidence measure exploits deep neural networks that are trained on large corpora without direct supervision. It is evaluated on an open-source spontaneous speech corpus and outperforms a confidence score derived from a state-of-the-art text-to-speech aligner. We further show that this confidence measure can be used to fine-tune the output of this aligner and improve the quality of the resulting alignment.
Code (0)
등록된 구현이 없습니다.
Tasks
speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Weakly-supervised forced alignment of disfluent speech using phoneme-level modeling
The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech aligners, hindering the use of automatic …
graph constructionGoodness-of-pronunciation without phoneme time alignment
In speech evaluation, an Automatic Speech Recognition (ASR) model often computes time boundaries and phoneme posteriors for input features. However, limited data for ASR training hinders expansion of speech evaluation to…
Speech RecognitionBitext Name Tagging for Cross-lingual Entity Annotation Projection
Annotation projection is a practical method to deal with the low resource problem in incident languages (IL) processing. Previous methods on annotation projection mainly relied on word alignment results without any train…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Leveraging Weakly Supervised Data to Improve End-to-End Speech-to-Text Translation
End-to-end Speech Translation (ST) models have many potential advantages when compared to the cascade of Automatic Speech Recognition (ASR) and text Machine Translation (MT) models, including lowered inference latency an…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationMulti-Task Learning+7Automatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech
In this paper, a framework for long audio alignment for conversational Arabic speech is proposed. Accurate alignments help in many speech processing tasks such as audio indexing, speech recognizer acoustic model (AM) tra…
Language Modellingspeech-recognitionSpeech Recognition