paper-with-me

홈 › Papers

Median-Based Generation of Synthetic Speech Durations using a Non-Parametric Approach

2016-08-22 · Srikanth Ronanki, Oliver Watts, Simon King, Gustav Eje Henter

This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike conventional approaches to duration modelling -- which assume that duration distributions have a particular form (e.g., a Gaussian) and use the mean of that distribution for synthesis -- our approach can in principle model any distribution supported on the non-negative integers. Generation from this model can be performed in many ways; here we consider output generation based on the median predicted duration. The median is more typical (more probable) than the conventional mean duration, is robust to training-data irregularities, and enables incremental generation. Furthermore, a frame-level approach to duration prediction is consistent with a longer-term goal of modelling durations and acoustic features together. Results indicate that the proposed method is competitive with baseline approaches in approximating the median duration of held-out natural speech.

📄 PDF Abstract BibTeX arXiv:1608.06134

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Neural Network-Based Modeling of Phonetic Durations

2019-09-06 · Xizi Wei, Melvyn Hunt, Adrian Skilling

A deep neural network (DNN)-based model has been developed to predict non-parametric distributions of durations of phonemes in specified phonetic contexts and used to explore which factors influence durations most. Major…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Expressive, Variable, and Controllable Duration Modelling in TTS

2022-06-28 · Ammar Abbas, Thomas Merritt, Alexis Moinet, Sri Karlapati 외

Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to relying on previous statistical parametric s…

Normalising FlowsSpeech Synthesistext-to-speechText to Speech

On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition

2023-10-12 · Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter

Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been shown that TTS-generated outputs still do n…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Parametric quantile autoregressive conditional duration models with application to intraday value-at-risk

2023-08-29 · Helton Saulo, Suvra Pal, Rubens Souza, Roberto Vila 외

The modeling of high-frequency data that qualify financial asset transactions has been an area of relevant interest among statisticians and econometricians -- above all, the analysis of time series of financial durations…

Diagnosticparameter estimation

Modeling the impact of control zone restrictions on pig placement in simulated African swine fever in the United States

2025-05-20 · Chunlin Yi, Jason A. Galvis, Gustavo Machado

African swine fever (ASF) is a highly contagious viral disease that poses a significant threat to the swine industry, requiring stringent control measures, including movement restrictions that delay pig placements, impac…