paper-with-me

홈 › Papers

TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models

2023-08-28 · Shengpeng Ji, Jialong Zuo, Minghui Fang, Ziyue Jiang, Feiyang Chen, Xinyu Duan, Baoxing Huai, Zhou Zhao

Recently, there has been a growing interest in the field of controllable Text-to-Speech (TTS). While previous studies have relied on users providing specific style factor values based on acoustic knowledge or selecting reference speeches that meet certain requirements, generating speech solely from natural text prompts has emerged as a new challenge for researchers. This challenge arises due to the scarcity of high-quality speech datasets with natural text style prompt and the absence of advanced text-controllable TTS models. In light of this, 1) we propose TextrolSpeech, which is the first large-scale speech emotion dataset annotated with rich text attributes. The dataset comprises 236,220 pairs of style prompt in natural text descriptions with five style factors and corresponding speech samples. Through iterative experimentation, we introduce a multi-stage prompt programming approach that effectively utilizes the GPT model for generating natural style descriptions in large volumes. 2) Furthermore, to address the need for generating audio with greater style diversity, we propose an efficient architecture called Salle. This architecture treats text controllable TTS as a language model task, utilizing audio codec codes as an intermediate representation to replace the conventional mel-spectrogram. Finally, we successfully demonstrate the ability of the proposed model by showing a comparable performance in the controllable TTS task. Audio samples are available at https://sall-e.github.io/

📄 PDF Abstract BibTeX arXiv:2308.14430

Code (1)

jishengpeng/TextrolSpeech 공식 구현

Tasks

Language Modellingtext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Weight Decay 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios

2021-12-23 · Qicong Xie, Tao Li, Xinsheng Wang, Zhichao Wang 외

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. …

DiversitySpeech SynthesisStyle Transfertext-to-speech+2

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

2018-03-23 · ICML 2018 7 · Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan 외

In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The embeddings are trained with no explicit lab…

Speech SynthesisStyle TransferText-To-Speech Synthesis

Expressive TTS Driven by Natural Language Prompts Using Few Human Annotations

2023-11-02 · Hanglei Zhang, Yiwei Guo, Sen Liu, Xie Chen 외

Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the ability to directly control synthesis st…

Language ModelingLanguage ModellingLarge Language ModelRetrieval+2

STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent

2022-03-28 · Yuki Saito, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana 외

We present STUDIES, a new speech corpus for developing a voice agent that can speak in a friendly manner. Humans naturally control their speech prosody to empathize with each other. By incorporating this "empathetic dial…

text-to-speechText to Speech

Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices

2022-06-01 · LREC 2022 6 · Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov 외

Synthetic voices are increasingly used in applications that require a conversational speaking style, raising the question as to which type of training data yields the most suitable speaking style for such applications. T…

Sentencetext-to-speechText to Speech