paper-with-me

홈 › Papers

The Theory behind Controllable Expressive Speech Synthesis: a Cross-disciplinary Approach

2019-10-14 · Noé Tits, Kevin El Haddad, Thierry Dutoit

As part of the Human-Computer Interaction field, Expressive speech synthesis is a very rich domain as it requires knowledge in areas such as machine learning, signal processing, sociology, psychology. In this Chapter, we will focus mostly on the technical side. From the recording of expressive speech to its modeling, the reader will have an overview of the main paradigms used in this field, through some of the most prominent systems and methods. We explain how speech can be represented and encoded with audio features. We present a history of the main methods of Text-to-Speech synthesis: concatenative, parametric and statistical parametric speech synthesis. Finally, we focus on the last one, with the last techniques modeling Text-to-Speech synthesis as a sequence-to-sequence problem. This enables the use of Deep Learning blocks such as Convolutional and Recurrent Neural Networks as well as Attention Mechanism. The last part of the Chapter intends to assemble the different aspects of the theory and summarize the concepts.

📄 PDF Abstract BibTeX arXiv:1910.06234

Code (0)

등록된 구현이 없습니다.

Tasks

Expressive Speech SynthesisSociologySpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech

2021-03-17 · Keon Lee, Kyumin Park, Daeyoung Kim

Previous works on neural text-to-speech (TTS) have been addressed on limited speed in training and inference time, robustness for difficult synthesis conditions, expressiveness, and controllability. Although several appr…

Speech SynthesisStyle Transfertext-to-speechText to Speech

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations

2025-05-24 · Ashwin Sankar, Yoach Lacombe, Sherry Thomas, Praveen Srinivasa Varadhan 외

We introduce RASMALAI, a large-scale speech dataset with rich text descriptions, designed to advance controllable and expressive text-to-speech (TTS) synthesis for 23 Indian languages and English. It comprises 13,000 hou…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech

Articulatory Phonetics Informed Controllable Expressive Speech Synthesis

2024-06-15 · Zehua Kcriss Li, Meiying Melissa Chen, Yi Zhong, Pinxin Liu 외

Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuan…

Expressive Speech SynthesisSpeech Synthesis

Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis

2025-12-21 · Pengchao Feng, Yao Xiao, Ziyang Ma, Zhikang Niu 외 arxiv

Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of ge…

Speech Synthesis

Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis through Audio Analysis

2019-03-27 · Noé Tits, Fengna Wang, Kevin El Haddad, Vincent Pagel 외

The field of Text-to-Speech has experienced huge improvements last years benefiting from deep learning techniques. Producing realistic speech becomes possible now. As a consequence, the research on the control of the exp…

Emotional Speech SynthesisExpressive Speech SynthesisLearning Network RepresentationsSpeech Emotion Recognition+4