paper-with-me

Papers

Unsupervised speech representation learning using WaveNet autoencoders

2019-01-25 · Jan Chorowski, Ron J. Weiss, Samy Bengio, Aäron van den Oord

We consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms. The goal is to learn a representation able to capture high level semantic content from the signal, e.g.\ phoneme identities, while being invariant to confounding low level details in the signal such as the underlying pitch contour or background noise. Since the learned representation is tuned to contain only phonetic content, we resort to using a high capacity WaveNet decoder to infer information discarded by the encoder from previous samples. Moreover, the behavior of autoencoder models depends on the kind of constraint that is applied to the latent representation. We compare three variants: a simple dimensionality reduction bottleneck, a Gaussian Variational Autoencoder (VAE), and a discrete Vector Quantized VAE (VQ-VAE). We analyze the quality of learned representations in terms of speaker independence, the ability to predict phonetic content, and the ability to accurately reconstruct individual spectrogram frames. Moreover, for discrete encodings extracted using the VQ-VAE, we measure the ease of mapping them to phonemes. We introduce a regularization scheme that forces the representations to focus on the phonetic content of the utterance and report performance comparable with the top entries in the ZeroSpeech 2017 unsupervised acoustic unit discovery task.

📄 PDF Abstract BibTeX arXiv:1901.08810

Code (5)

MingjieChen/wavenet_autoencoders pytorch
StanislavParovoy/VQ-VAE-WaveNet tf
bshall/ZeroSpeech pytorch
hrbigelow/ae-wavenet pytorch
swasun/VQ-VAE-Speech pytorch

Tasks

Acoustic Unit DiscoveryDecoderDimensionality ReductionRepresentation LearningSpeech Representation Learning

Methods 이 논문이 사용한 방법론

Mixture of Logistic Distributions 설명 없음
VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…
Solana Customer Service Number +1-833-534-1729 설명 없음
Dilated Causal Convolution A Dilated Causal Convolution is a causal convolution where the filter is applied over an area larger than its length by…
WaveNet WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies…
USD Coin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Do WaveNets Dream of Acoustic Waves?

2018-02-23 · Kanru Hua

Various sources have reported the WaveNet deep learning architecture being able to generate high-quality speech, but to our knowledge there haven't been studies on the interpretation or visualization of trained WaveNets.…

Speech Representations and Phoneme Classification for Preserving the Endangered Language of Ladin

2021-08-27 · Zane Durante, Leena Mathur, Eric Ye, Sichong Zhao 외

A vast majority of the world's 7,000 spoken languages are predicted to become extinct within this century, including the endangered language of Ladin from the Italian Alps. Linguists who work to preserve a language's pho…

Parallel WaveNet conditioned on VAE latent vectors

2020-12-17 · Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba, Adam Gabrys 외

Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically mel-spectrograms), followed by a 'neural …

SentenceSpeech Synthesistext-to-speechText to Speech+1

Multi-task WaveNet: A Multi-task Generative Model for Statistical Parametric Speech Synthesis without Fundamental Frequency Conditions

2018-06-22

This paper introduces an improved generative model for statistical parametric speech synthesis (SPSS) based on WaveNet under a multi-task learning framework. Different from the original WaveNet model, the proposed Multi-…

Multi-Task LearningPredictionSpeech Synthesis

Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

2017-12-16 · Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster 외

This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings…

Speech Synthesis