paper-with-me

홈 › Papers

A Fully Time-domain Neural Model for Subband-based Speech Synthesizer

2018-10-22 · Anonymous

—This paper introduces a deep neural network model for subband-based speech synthesizer. The model benefits from the short bandwidth of the subband signals to reduce the complexity of the time-domain speech generator. We employed the multi-level wavelet analysis/synthesis to decompose/reconstruct the signal to subbands in time domain. Inspired from the WaveNet, a convolutional neural network (CNN) model predicts subband speech signals fully in time domain. Due to the short bandwidth of the subbands, a simple network architecture is enough to train the simple patterns of the subbands accurately. In the ground truth experiments with teacher forcing, the subband synthesizer outperforms the fullband model significantly. In addition, by conditioning the model on the phoneme sequence using a pronunciation dictionary, we have achieved the first fully time-domain neural text-to-speech (TTS) system. The generated speech of the subband TTS shows comparable quality as the fullband one with a slighter network architecture for each subband.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

A Fully Time-domain Neural Model for Subband-based Speech Synthesizer

2018-10-12 · Azam Rabiee, Soo-Young Lee

This paper introduces a deep neural network model for subband-based speech synthesizer. The model benefits from the short bandwidth of the subband signals to reduce the complexity of the time-domain speech generator. We …

text-to-speechText to Speech

Extending DNN-based Multiplicative Masking to Deep Subband Filtering for Improved Dereverberation

2023-03-01 · Jean-Marie Lemercier, Julian Tobergte, Timo Gerkmann

In this paper, we present a scheme for extending deep neural network-based multiplicative maskers to deep subband filters for speech restoration in the time-frequency domain. The resulting method can be generically appli…

Denoising

Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario

2022-10-14 · Emily R. Bartusiak, Edward J. Delp

Speech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and misinformation campaigns. Forensic methods that detect synthesized speech are important for protection against suc…

AttributeMisinformationMulti-class ClassificationSpeech Synthesis

Transferring neural speech waveform synthesizers to musical instrument sounds generation

2019-10-27 · Yi Zhao, Xin Wang, Lauri Juvela, Junichi Yamagishi

Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similari…

Audio GenerationAudio SynthesisSpeech SynthesisZero-Shot Learning

Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation

2024-06-14 · Nameer Hirschkind, Xiao Yu, Mahesh Kumar Nandwana, Joseph Liu 외

We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment wit…

Speech-to-Speech TranslationTranslation