paper-with-me

Papers

LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

2024-06-12 · Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi, Kentaro Tachibana

We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to construct prompt annotations: (1) manual annotations that capture human perceptions of speaker characteristics and (2) synthetic annotations on speaking style. Compared to existing English prompt datasets, our corpus provides more diverse prompt annotations for all speakers of LibriTTS-R. Experimental results for prompt-based controllable TTS demonstrate that the TTS model trained with LibriTTS-P achieves higher naturalness than the model using the conventional dataset. Furthermore, the results for style captioning tasks show that the model utilizing LibriTTS-P generates 2.5 times more accurate words than the model using a conventional dataset. Our corpus, LibriTTS-P, is available at https://github.com/line/LibriTTS-P.

📄 PDF Abstract BibTeX arXiv:2406.07969

Code (1)

line/libritts-p 공식 구현

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions

2023-09-15 · Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata 외

We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framewor…

text-to-speechText to Speech

LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

2023-05-30 · Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 외

This paper introduces a new speech dataset called ``LibriTTS-R'' designed for text-to-speech (TTS) use. It is derived by applying speech restoration to the LibriTTS corpus, which consists of 585 hours of speech data at 2…

text-to-speechText to Speech

Cross-speaker style transfer for text-to-speech using data augmentation

2022-02-10 · Manuel Sam Ribeiro, Julian Roth, Giulia Comini, Goeric Huybrechts 외

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting…

Data AugmentationStyle Transfertext-to-speechText to Speech+1

Imaginary Voice: Face-styled Diffusion Model for Text-to-Speech

2023-02-27 · Jiyoung Lee, Joon Son Chung, Soo-Whan Chung

The goal of this work is zero-shot text-to-speech synthesis, with speaking styles and voices learnt from facial characteristics. Inspired by the natural fact that people can imagine the voice of someone when they look at…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

2018-03-23 · ICML 2018 7 · Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan 외

In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The embeddings are trained with no explicit lab…

Speech SynthesisStyle TransferText-To-Speech Synthesis