paper-with-me

Papers

Into-TTS : Intonation Template Based Prosody Control System

2022-04-04 · JIhwan Lee, Joun Yeop Lee, Heejin Choi, Seongkyu Mun, Sangjun Park, Jae-Sung Bae, Chanwoo Kim

Intonations play an important role in delivering the intention of a speaker. However, current end-to-end TTS systems often fail to model proper intonations. To alleviate this problem, we propose a novel, intuitive method to synthesize speech in different intonations using predefined intonation templates. Prior to TTS model training, speech data are grouped into intonation templates in an unsupervised manner. Two proposed modules are added to the end-to-end TTS framework: an intonation predictor and an intonation encoder. The intonation predictor recommends a suitable intonation template to the given text. The intonation encoder, attached to the text encoder output, synthesizes speech abiding the requested intonation template. Main contributions of our paper are: (a) an easy-to-use intonation control system covering a wide range of users; (b) better performance in wrapping speech in a requested intonation with improved objective and subjective evaluation; and (c) incorporating a pre-trained language model for intonation modelling. Audio samples are available at https://srtts.github.io/IntoTTS.

📄 PDF Abstract BibTeX arXiv:2204.01271

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis

2024-12-16 · Xiangheng He, Junjie Chen, Zixing Zhang, Björn W. Schuller

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of F0

2020-03-14 · Zack Hodari, Catherine Lai, Simon King

In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However, when learning to control prosody in tex…

ClusteringRepresentation LearningSpeech Synthesistext-to-speech+1

Using generative modelling to produce varied intonation for speech synthesis

2019-06-10 · Zack Hodari, Oliver Watts, Simon King

Unlike human speakers, typical text-to-speech (TTS) systems are unable to produce multiple distinct renditions of a given sentence. This has previously been addressed by adding explicit external control. In contrast, gen…

SentenceSpeech Synthesistext-to-speechText to Speech

Word-wise intonation model for cross-language TTS systems

2024-09-30 · Tomilov A. A., Gromova A. Y., Svischev A. N

In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to …

Dynamic Time WarpingProsody Predictiontext-to-speechText to Speech

Prosody Labelled Dataset for Hindi

2021-12-01 · SMP (ICON) 2021 12 · Esha Banerjee, Atul Kr. Ojha, Girish Jha

This study aims to develop an intonation labelled database for Hindi, for enhancing prosody in ASR and TTS systems, which is also helpful for building Speech to Speech Machine Translation systems. Although no single stan…

Machine TranslationTranslation