PITS: Variational Pitch Inference without Fundamental Frequency for End-to-End Pitch-controllable TTS
Previous pitch-controllable text-to-speech (TTS) models rely on directly modeling fundamental frequency, leading to low variance in synthesized speech. To address this issue, we propose PITS, an end-to-end pitch-controllable TTS model that utilizes variational inference to model pitch. Based on VITS, PITS incorporates the Yingram encoder, the Yingram decoder, and adversarial training of pitch-shifted synthesis to achieve pitch-controllability. Experiments demonstrate that PITS generates high-quality speech that is indistinguishable from ground truth speech and has high pitch-controllability without quality degradation. Code, audio samples, and demo are available at https://github.com/anonymous-pits/pits.
Code (2)
Tasks
Decodertext-to-speechText to SpeechVariational InferenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Variational Inference of Structured Line Spectra Exploiting Group-Sparsity
In this paper, we present a variational inference algorithm that decomposes a signal into multiple groups of related spectral lines. The spectral lines in each group are associated with a group parameter common to all sp…
Variational InferencePeriod VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis
Several fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate …
DecoderDiversityEmotional Speech SynthesisSpeech Synthesis+3Cross-modal variational inference for bijective signal-symbol translation
Extraction of symbolic information from signals is an active field of research enabling numerous applications especially in the Musical Information Retrieval domain. This complex task, that is also related to other topic…
Audio GenerationDensity EstimationInformation RetrievalInstrument Recognition+4FastPitch: Parallel Text-to-speech with Pitch Prediction
We present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours. The model predicts pitch contours during inference. By altering these predictions, the gener…
Predictiontext-to-speechText to SpeechPitch and timbre discrimination at wave-to-spike transition in the cochlea
A new definition of musical pitch is proposed. A Finite-Difference Time Domain (FDTM) model of the cochlea is used to calculate spike trains caused by tone complexes and by a recorded classical guitar tone. All harmonic …