paper-with-me

Papers

On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition

2023-10-12 · Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter

Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been shown that TTS-generated outputs still do not have the same qualities as real data. In this work we focus on the temporal structure of synthetic data and its relation to ASR training. By using a novel oracle setup we show how much the degradation of synthetic data quality is influenced by duration modeling in non-autoregressive (NAR) TTS. To get reference phoneme durations we use two common alignment methods, a hidden Markov Gaussian-mixture model (HMM-GMM) aligner and a neural connectionist temporal classification (CTC) aligner. Using a simple algorithm based on random walks we shift phoneme duration distributions of the TTS system closer to real durations, resulting in an improvement of an ASR system using synthetic data in a semi-supervised setting.

📄 PDF Abstract BibTeX arXiv:2310.08132

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognitiontext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Reinforce-Aligner: Reinforcement Alignment Search for Robust End-to-End Text-to-Speech

2021-06-05 · Hyunseung Chung, Sang-Hoon Lee, Seong-Whan Lee

Text-to-speech (TTS) synthesis is the process of producing synthesized speech from text or phoneme input. Traditional TTS models contain multiple processing steps and require external aligners, which provide attention al…

text-to-speechText to Speech

Phone Duration Modeling for Speaker Age Estimation in Children

2021-09-03 · Prashanth Gurunath Shivakumar, Somer Bishop, Catherine Lord, Shrikanth Narayanan

Automatic inference of important paralinguistic information such as age from speech is an important area of research with numerous spoken language technology based applications. Speaker age estimation has applications in…

Age Estimationregression

Controllable speech synthesis by learning discrete phoneme-level prosodic representations

2022-11-29 · Nikolaos Ellinas, Myrsini Christidou, Alexandra Vioni, June Sig Sung 외

In this paper, we present a novel method for phoneme-level prosody control of F0 and duration using intuitive discrete labels. We propose an unsupervised prosodic clustering process which is used to discretize phoneme-le…

ClusteringSpeech Synthesistext-to-speechText to Speech

Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control

2021-11-19 · Myrsini Christidou, Alexandra Vioni, Nikolaos Ellinas, Georgios Vamvoukakis 외

This paper presents a method for phoneme-level prosody control of F0 and duration on a multispeaker text-to-speech setup, which is based on prosodic clustering. An autoregressive attention-based model is used, incorporat…

ClusteringData Augmentationtext-to-speechText to Speech

JDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment

2020-05-15 · Dan Lim, Won Jang, Gyeonghwan O, Heayoung Park 외

We propose Jointly trained Duration Informed Transformer (JDI-T), a feed-forward Transformer with a duration predictor jointly trained without explicit alignments in order to generate an acoustic feature sequence from an…

text-to-speechText to Speech