paper-with-me

Papers

Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis

2018-08-04 · Daisy Stanton, Yuxuan Wang, RJ Skerry-Ryan

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis system, to uncover expressive factors of variation in speaking style. In this work, we introduce the Text-Predicted Global Style Token (TP-GST) architecture, which treats GST combination weights or style embeddings as "virtual" speaking style labels within Tacotron. TP-GST learns to predict stylistic renderings from text alone, requiring neither explicit labels during training nor auxiliary inputs for inference. We show that, when trained on a dataset of expressive speech, our system generates audio with more pitch and energy variation than two state-of-the-art baseline models. We further demonstrate that TP-GSTs can synthesize speech with background noise removed, and corroborate these analyses with positive results on human-rated listener preference audiobook tasks. Finally, we demonstrate that multi-speaker TP-GST models successfully factorize speaker identity and speaking style. We provide a website with audio samples for each of our findings.

📄 PDF Abstract BibTeX arXiv:1808.01410

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Griffin-Lim Algorithm The Griffin-Lim Algorithm (GLA) is a phase reconstruction method based on the redundancy of the short-time Fourier transform. It promotes the consistency of a spectrogram by…
Sigmoid Activation 설명 없음
Highway Layer 설명 없음
Residual Connection 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Batch Normalization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Residual GRU A Residual GRU is a gated recurrent unit (GRU) that incorporates the idea of residual connections from…

Similar Papers 제목 키워드 기반

Expressive Text-to-Speech using Style Tag

2021-04-01 · Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim 외

As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expre…

Language ModelingLanguage ModellingTAGtext-to-speech+1

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

2025-08-15 · Wonjune Kang, Deb Roy arxiv

We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused …

Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

2022-03-23 · Shun Lei, Yixuan Zhou, Liyang Chen, Zhiyong Wu 외

Previous works on expressive speech synthesis mainly focus on current sentence. The context in adjacent sentences is neglected, resulting in inflexible speaking style for the same text, which lacks speech variations. In …

Expressive Speech SynthesisKnowledge DistillationSentenceSpeech Synthesis

Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis

2022-10-01 · COLING 2022 10 · Xueyuan Chen, Shun Lei, Zhiyong Wu, Dong Xu 외

Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sen…

Speech Synthesistext-to-speechText to Speech

Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis

2021-09-08 · Songxiang Liu, Shan Yang, Dan Su, Dong Yu

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSST approaches rely on expensive high-qual…

Expressive Speech SynthesisSentenceSpeech SynthesisStyle Transfer+2