paper-with-me

Papers

In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data

2019-04-04 · NAACL 2019 6 · Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote, Thomas Drugman, Jaime Lorenzo-Trueba, Thomas Merritt, Srikanth Ronanki, Trevor Wood

Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes creating models for multiple styles expensive and time-consuming. In this paper different styles of speech are analysed based on prosodic variations, from this a model is proposed to synthesise speech in the style of a newscaster, with just a few hours of supplementary data. We pose the problem of synthesising in a target style using limited data as that of creating a bi-style model that can synthesise both neutral-style and newscaster-style speech via a one-hot vector which factorises the two styles. We also propose conditioning the model on contextual word embeddings, and extensively evaluate it against neutral NTTS, and neutral concatenative-based synthesis. This model closes the gap in perceived style-appropriateness between natural recordings for newscaster-style of speech, and neutral speech synthesis by approximately two-thirds.

📄 PDF Abstract BibTeX arXiv:1904.02790

Code (1)

inconnu11/Objective-evaluation_speech_synthesis

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisWord Embeddings

Similar Papers 제목 키워드 기반

Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM

2024-11-20 · Jiawei Yu, Yuang Li, Xiaosong Qiao, Huan Zhao 외

Text-to-speech (TTS) models have been widely adopted to enhance automatic speech recognition (ASR) systems using text-only corpora, thereby reducing the cost of labeling real speech data. Existing research primarily util…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

AdaSpeech 3: Adaptive Text to Speech for Spontaneous Style

2021-07-06 · Yuzi Yan, Xu Tan, Bohan Li, Guangyan Zhang 외

While recent text to speech (TTS) models perform very well in synthesizing reading-style (e.g., audiobook) speech, it is still challenging to synthesize spontaneous-style speech (e.g., podcast or conversation), mainly be…

DecoderMixture-of-ExpertsRhythmtext-to-speech+1

Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

2025-11-18 · Nam-Gyu Kim arxiv

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We pr…

Style Transfer

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

2025-05-27 · Nam-Gyu Kim, Deok-Hyeon Cho, Seung-bin Kim, Seong-Whan Lee

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We pr…

Style Transfertext-to-speechText to Speech

Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis

2023-08-31 · Weiqin Li, Shun Lei, Qiaochu Huang, Yixuan Zhou 외

The spontaneous behavior that often occurs in conversations makes speech more human-like compared to reading-style. However, synthesizing spontaneous-style speech is challenging due to the lack of high-quality spontaneou…

Expressive Speech SynthesisSentenceSpeech Synthesistext-to-speech+2