paper-with-me

홈 › Papers

Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch

2024-10-09 · Teodora Răgman, Adriana Stan

This paper focuses on adapting the functionalities of the FastPitch model to the Romanian language; extending the set of speakers from one to eighteen; synthesising speech using an anonymous identity; and replicating the identities of new, unseen speakers. During this work, the effects of various configurations and training strategies were tested and discussed, along with their advantages and weaknesses. Finally, we settled on a new configuration, built on top of the FastPitch architecture, capable of producing natural speech synthesis, for both known (identities from the training dataset) and unknown (identities learnt through short reference samples) speakers. The anonymous speaker can be used for text-to-speech synthesis, if one wants to cancel out the identity information while keeping the semantic content whole and clear. At last, we discussed possible limitations of our work, which will form the basis for future investigations and advancements.

📄 PDF Abstract BibTeX arXiv:2410.06787

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Attention 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Auto Spell Suggestion for High Quality Speech Synthesis in Hindi

2014-02-15 · Shikha Kabra, Ritika Agarwal

The goal of Text-to-Speech (TTS) synthesis in a particular language is to convert arbitrary input text to intelligible and natural sounding speech. However, for a particular language like Hindi, which is a highly confusi…

Speech Synthesistext-to-speechText to SpeechVocal Bursts Intensity Prediction

LoRP-TTS: Low-Rank Personalized Text-To-Speech

2025-02-11 · Łukasz Bondaruk, Jakub Kubiak

Speech synthesis models convert written text into natural-sounding audio. While earlier models were limited to a single speaker, recent advancements have led to the development of zero-shot systems that generate realisti…

Speech Synthesistext-to-speechText to Speech

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

2024-12-09 · Tianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 외

Text-to-speech (TTS), also known as speech synthesis, is a prominent research area that aims to generate natural-sounding human speech from text. Recently, with the increasing industrial demand, TTS technologies have evo…

Speech SynthesisSurveytext-to-speechText to Speech

Speech Recognition with Augmented Synthesized Speech

2019-09-25 · Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia 외

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribe…

Data AugmentationDiversityRobust Speech Recognitionspeech-recognition+2

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

2026-06-12 · Vinh Dang Quang, Huy Ngo Quang arxiv

For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep learning-based TTS systems developed to make it possible to produce voices with …

Speech Synthesis