paper-with-me

Papers

ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading

2023-07-03 · Yujia Xiao, Shaofei Zhang, Xi Wang, Xu Tan, Lei He, Sheng Zhao, Frank K. Soong, Tan Lee

While state-of-the-art Text-to-Speech systems can generate natural speech of very high quality at sentence level, they still meet great challenges in speech generation for paragraph / long-form reading. Such deficiencies are due to i) ignorance of cross-sentence contextual information, and ii) high computation and memory cost for long-form synthesis. To address these issues, this work develops a lightweight yet effective TTS system, ContextSpeech. Specifically, we first design a memory-cached recurrence mechanism to incorporate global text and speech context into sentence encoding. Then we construct hierarchically-structured textual semantics to broaden the scope for global context enhancement. Additionally, we integrate linearized self-attention to improve model efficiency. Experiments show that ContextSpeech significantly improves the voice quality and prosody expressiveness in paragraph reading with competitive model efficiency. Audio samples are available at: https://contextspeech.github.io/demo/

📄 PDF Abstract BibTeX arXiv:2307.00782

Code (0)

등록된 구현이 없습니다.

Tasks

FormSentencetext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis

2020-11-06 · Guanghui Xu, Wei Song, Zhengchen Zhang, Chao Zhang 외

Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which makes it challenging when converting a par…

DecoderSentenceSentence EmbeddingsSpeech Synthesis+2

Discovering the Italian literature: interactive access to audio indexed text resources

2014-05-01 · LREC 2014 5 · Vincenzo Galat{\`a}, Alberto Benin, Piero Cosi, Giuseppe Riccardo Leone 외

In this paper we present a web interface to study Italian through the access to read Italian literature. The system allows to browse the content, search for specific words and listen to the correct pronunciation produced…

Cultural Vocal Bursts Intensity PredictionSentencetext-to-speechText to Speech

Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

2022-06-25 · Yihan Wu, Xi Wang, Shaofei Zhang, Lei He 외

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of lab…

Contrastive LearningDeep ClusteringExpressive Speech SynthesisRepresentation Learning+1

How to Train Your Agent to Read and Write

2021-01-04 · Li Liu, Mengge He, Guanghui Xu, Mingkui Tan 외

Reading and writing research papers is one of the most privileged abilities that a qualified researcher should master. However, it is difficult for new researchers (\eg{students}) to fully {grasp} this ability. It would …

KG-to-Text GenerationKnowledge Graphs

STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech

2021-03-17 · Keon Lee, Kyumin Park, Daeyoung Kim

Previous works on neural text-to-speech (TTS) have been addressed on limited speed in training and inference time, robustness for difficult synthesis conditions, expressiveness, and controllability. Although several appr…

Speech SynthesisStyle Transfertext-to-speechText to Speech