paper-with-me

Papers

Simple and Effective Unsupervised Speech Synthesis

2022-04-06 · Alexander H. Liu, Cheng-I Jeff Lai, Wei-Ning Hsu, Michael Auli, Alexei Baevski, James Glass

We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe. The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesis. Using only unlabeled speech audio and unlabeled text as well as a lexicon, our method enables speech synthesis without the need for a human-labeled corpus. Experiments demonstrate the unsupervised system can synthesize speech similar to a supervised counterpart in terms of naturalness and intelligibility measured by human evaluation.

📄 PDF Abstract BibTeX arXiv:2204.02524

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech SynthesisUnsupervised Speech Recognition

Similar Papers 제목 키워드 기반

Simple and Effective Unsupervised Speech Translation

2022-10-18 · Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov 외

The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled data covering two different languages. To…

Domain AdaptationMachine Translationspeech-recognitionSpeech Recognition+6

Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition

2022-03-29 · Junrui Ni, Liming Wang, Heting Gao, Kaizhi Qian 외

An unsupervised text-to-speech synthesis (TTS) system learns to generate speech waveforms corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that langua…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+5

Unsupervised Audiovisual Synthesis via Exemplar Autoencoders

2020-01-13 · ICLR 2021 1 · Kangle Deng, Aayush Bansal, Deva Ramanan

We present an unsupervised approach that converts the input speech of any individual into audiovisual streams of potentially-infinitely many output speakers. Our approach builds on simple autoencoders that project out-of…

Multimodal speech synthesis architecture for unsupervised speaker adaptation

2018-08-20 · Hieu-Thi Luong, Junichi Yamagishi

This paper proposes a new architecture for speaker adaptation of multi-speaker neural-network speech synthesis systems, in which an unseen speaker's voice can be built using a relatively small amount of speech data witho…

Speech Synthesis

Expediting TTS Synthesis with Adversarial Vocoding

2019-04-16 · Paarth Neekhara, Chris Donahue, Miller Puckette, Shlomo Dubnov 외

Recent approaches in text-to-speech (TTS) synthesis employ neural network strategies to vocode perceptually-informed spectrogram representations directly into listenable waveforms. Such vocoding procedures create a compu…

text-to-speechText to Speech