paper-with-me

홈 › Papers

Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder

2018-04-06 · Kei Akuzawa, Yusuke Iwasawa, Yutaka Matsuo

Recent advances in neural autoregressive models have improve the performance of speech synthesis (SS). However, as they lack the ability to model global characteristics of speech (such as speaker individualities or speaking styles), particularly when these characteristics have not been labeled, making neural autoregressive SS systems more expressive is still an open issue. In this paper, we propose to combine VoiceLoop, an autoregressive SS model, with Variational Autoencoder (VAE). This approach, unlike traditional autoregressive SS systems, uses VAE to model the global characteristics explicitly, enabling the expressiveness of the synthesized speech to be controlled in an unsupervised manner. Experiments using the VCTK and Blizzard2012 datasets show the VAE helps VoiceLoop to generate higher quality speech and to control the expressions in its synthesized speech by incorporating global characteristics into the speech generating process.

📄 PDF Abstract BibTeX arXiv:1804.02135

Code (0)

등록된 구현이 없습니다.

Tasks

Expressive Speech SynthesisSpeech Synthesis

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음
USD Coin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

2025-03-13 · Yasheng Sun, Zhiliang Xu, Hang Zhou, Jiazhi Guan 외

Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these c…

DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis

2024-10-17 · Yu Gu, Qiushi Zhu, Guangzhi Lei, Chao Weng 외

This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E mode…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Laughter Synthesis: Combining Seq2seq modeling with Transfer Learning

2020-08-20 · Noé Tits, Kevin El Haddad, Thierry Dutoit

Despite the growing interest for expressive speech synthesis, synthesis of nonverbal expressions is an under-explored area. In this paper we propose an audio laughter synthesis system based on a sequence-to-sequence TTS …

Expressive Speech SynthesisSpeech SynthesisTransfer Learning

MIBURI: Towards Expressive Interactive Gesture Synthesis

2026-03-03 · M. Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian Theobalt arxiv

Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conversational agents lack embodiment and the…

Fine-grained Noise Control for Multispeaker Speech Synthesis

2022-04-11 · Karolos Nikitaras, Georgios Vamvoukakis, Nikolaos Ellinas, Konstantinos Klapsas 외

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in ord…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech