paper-with-me

홈 › Papers

Controlling Prosody in End-to-End TTS: A Case Study on Contrastive Focus Generation

2021-11-01 · CoNLL (EMNLP) 2021 11 · Siddique Latif, Inyoung Kim, Ioan Calapodescu, Laurent Besacier

While End-2-End Text-to-Speech (TTS) has made significant progresses over the past few years, these systems still lack intuitive user controls over prosody. For instance, generating speech with fine-grained prosody control (prosodic prominence, contextually appropriate emotions) is still an open challenge. In this paper, we investigate whether we can control prosody directly from the input text, in order to code information related to contrastive focus which emphasizes a specific word that is contrary to the presuppositions of the interlocutor. We build and share a specific dataset for this purpose and show that it allows to train a TTS system were this fine-grained prosodic feature can be correctly conveyed using control tokens. Our evaluation compares synthetic and natural utterances and shows that prosodic patterns of contrastive focus (variations of Fo, Intensity and Duration) can be learnt accurately. Such a milestone is important to allow, for example, smart speakers to be programmatically controlled in terms of output prosody.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model

2022-07-04 · Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber

Several recent studies have tested the use of transformer language model representations to infer prosodic features for text-to-speech synthesis (TTS). While these studies have explored prosody in general, in this work, …

Language ModelingLanguage ModellingSpeech Synthesistext-to-speech+2

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

2025-07-27 · Kaizhi Qian, Xulin Fan, Junrui Ni, Slava Shechtman 외 arxiv

Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to capture the intricate interdependency betwe…

Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases

2024-02-01 · Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow

Speech-to-Text Translation (S2TT) has typically been addressed with cascade systems, where speech recognition systems generate a transcription that is subsequently passed to a translation model. While there has been a gr…

speech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text Translation+1

Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?

2024-10-31 · Ioannis Tsiamas, Matthias Sperber, Andrew Finch, Sarthak Garg

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, pro…

Rhythmspeech-recognitionSpeech RecognitionSpeech-to-Text+4

Cloning one's voice using very limited data in the wild

2021-10-07 · Dongyang Dai, Yuanzhe Chen, Li Chen, Ming Tu 외

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily accessible data to clone a person's voice…

Speech Synthesis