paper-with-me

Papers

Noise Robust TTS for Low Resource Speakers using Pre-trained Model and Speech Enhancement

2020-05-26 · Dongyang Dai, Li Chen, Yu-Ping Wang, Mu Wang, Rui Xia, Xuchen Song, Zhiyong Wu, Yuxuan Wang

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech synthesis technology have been widely used in our daily life. Robust speech synthesis model depends on high quality and customized data which needs lots of collecting efforts. It is worth investigating how to take advantage of low-quality and low resource voice data which can be easily obtained from the Internet for usage of synthesizing personalized voice. In this paper, the proposed end-to-end speech synthesis model uses both speaker embedding and noise representation as conditional inputs to model speaker and noise information respectively. Firstly, the speech synthesis model is pre-trained with both multi-speaker clean data and noisy augmented data; then the pre-trained model is adapted on noisy low-resource new speaker data; finally, by setting the clean speech condition, the model can synthesize the new speaker's clean voice. Experimental results show that the speech generated by the proposed approach has better subjective evaluation results than the method directly fine-tuning pre-trained multi-speaker speech synthesis model with denoised new speaker data.

📄 PDF Abstract BibTeX arXiv:2005.12531

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech EnhancementSpeech Synthesis

Similar Papers 제목 키워드 기반

Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement

2019-11-14 · Soumi Maiti, Michael I Mandel

Traditional speech enhancement systems produce speech with compromised quality. Here we propose to use the high quality speech generation capability of neural vocoders for better quality speech enhancement. We term this …

ResynthesisSpeech Enhancement

Speaker Re-identification with Speaker Dependent Speech Enhancement

2020-05-15 · Yanpei Shi, Qiang Huang, Thomas Hain

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditiona…

Speaker RecognitionSpeech Enhancement

Real-Time System for Audio-Visual Target Speech Enhancement

2025-09-25 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…

Audio-Visual Speech RecognitionSpeech Enhancement

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

2025-07-29 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…

Audio-Visual Speech RecognitionActive Speaker DetectionSpeech Enhancement

TAPS: Throat and Acoustic Paired Speech Dataset for Deep Learning-Based Speech Enhancement

2025-02-17 · Yunsik Kim, Yonghun Song, Yoonyoung Chung

In high-noise environments such as factories, subways, and busy streets, capturing clear speech is challenging. Throat microphones can offer a solution because of their inherent noise-suppression capabilities; however, t…

Speech Enhancement