paper-with-me

Papers

Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis

2022-10-01 · COLING 2022 10 · Xueyuan Chen, Shun Lei, Zhiyong Wu, Dong Xu, Weifeng Zhao, Helen Meng

Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks. A multi-scale hierarchical context encoder is specially designed to predict both global-scale context style embedding and local-scale context style embedding from a wider context of input text in a hierarchical manner. Likewise, a multi-scale reference encoder is introduced to extract reference style embeddings at both global and local scales from the reference speech, which is used to guide the prediction of speaking styles. On top of these, a bi-reference attention mechanism is used to align both local-scale reference style embedding sequence and local-scale context style embedding sequence with corresponding phoneme embedding sequence. Both objective and subjective experiment results on a real-world multi-speaker Mandarin novel audio dataset demonstrate the excellent performance of our proposed method over all baselines in terms of naturalness and expressiveness of the synthesized speech.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling

2023-09-03 · Zhichao Wang, Xinsheng Wang, Qicong Xie, Tao Li 외

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice conversion (VC) task, which is essential i…

Data AugmentationDisentanglementEmotion RecognitionSpeech Emotion Recognition+2

Expressive Text-to-Speech using Style Tag

2021-04-01 · Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim 외

As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expre…

Language ModelingLanguage ModellingTAGtext-to-speech+1

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

2018-03-23 · ICML 2018 7 · Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan 외

In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The embeddings are trained with no explicit lab…

Speech SynthesisStyle TransferText-To-Speech Synthesis

Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis

2018-08-04 · Daisy Stanton, Yuxuan Wang, RJ Skerry-Ryan

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis sy…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

2025-08-15 · Wonjune Kang, Deb Roy arxiv

We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused …