paper-with-me

Papers

AS-Speech: Adaptive Style For Speech Synthesis

2024-09-09 · Zhipeng Li, Xiaofen Xing, Jun Wang, Shuaiqi Chen, Guoqiao Yu, Guanglu Wan, Xiangmin Xu

In recent years, there has been significant progress in Text-to-Speech (TTS) synthesis technology, enabling the high-quality synthesis of voices in common scenarios. In unseen situations, adaptive TTS requires a strong generalization capability to speaker style characteristics. However, the existing adaptive methods can only extract and integrate coarse-grained timbre or mixed rhythm attributes separately. In this paper, we propose AS-Speech, an adaptive style methodology that integrates the speaker timbre characteristics and rhythmic attributes into a unified framework for text-to-speech synthesis. Specifically, AS-Speech can accurately simulate style characteristics through fine-grained text-based timbre features and global rhythm information, and achieve high-fidelity speech synthesis through the diffusion model. Experiments show that the proposed model produces voices with higher naturalness and similarity in terms of timbre and rhythm compared to a series of adaptive TTS models.

📄 PDF Abstract BibTeX arXiv:2409.05730

Code (0)

등록된 구현이 없습니다.

Tasks

RhythmSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models

2022-11-17 · Minki Kang, Dongchan Min, Sung Ju Hwang

There has been a significant progress in Text-To-Speech (TTS) synthesis technology in recent years, thanks to the advancement in neural generative modeling. However, existing methods on any-speaker adaptive TTS have achi…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Few Shot Adaptive Normalization Driven Multi-Speaker Speech Synthesis

2020-12-14 · Neeraj Kumar, Srishti Goel, Ankur Narang, Brejesh lall

The style of the speech varies from person to person and every person exhibits his or her own style of speaking that is determined by the language, geography, culture and other factors. Style is best captured by prosody …

Cultural Vocal Bursts Intensity PredictionSpeech Synthesis

StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization

2020-11-03 · Ahmed Mustafa, Nicola Pia, Guillaume Fuchs

In recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best …

Spectral Reconstructiontext-to-speechText to SpeechVocal Bursts Intensity Prediction

In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data

2019-04-04 · NAACL 2019 6 · Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote, Thomas Drugman 외

Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes creating models for multiple styles expe…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1

Towards Multi-Scale Style Control for Expressive Speech Synthesis

2021-04-08 · Xiang Li, Changhe Song, Jingbei Li, Zhiyong Wu 외

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level an…

Expressive Speech SynthesisSpeech SynthesisStyle Transfer