paper-with-me

홈 › Papers

A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning

2020-06-03 · Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski, Adrian Lancucki, Ricard Marxer, James Glass

Probabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the latent structure shapes the information extracted from the signal. Even though LVMs have recently seen a renewed interest due to the introduction of Variational Autoencoders (VAEs), their use for speech representation learning remains largely unexplored. In this work, we propose Convolutional Deep Markov Model (ConvDMM), a Gaussian state-space model with non-linear emission and transition functions modelled by deep neural networks. This unsupervised model is trained using black box variational inference. A deep convolutional neural network is used as an inference network for structured variational approximation. When trained on a large scale speech dataset (LibriSpeech), ConvDMM produces features that significantly outperform multiple self-supervised feature extracting methods on linear phone classification and recognition on the Wall Street Journal dataset. Furthermore, we found that ConvDMM complements self-supervised methods like Wav2Vec and PASE, improving on the results achieved with any of the methods alone. Lastly, we find that ConvDMM features enable learning better phone recognizers than any other features in an extreme low-resource regime with few labeled training examples.

📄 PDF Abstract BibTeX arXiv:2006.02547

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSelf-Supervised LearningSpeech Representation LearningVariational Inference

Similar Papers 제목 키워드 기반

Bio-Inspired Multi-Layer Spiking Neural Network Extracts Discriminative Features from Speech Signals

2017-06-10 · Amirhossein Tavanaei, Anthony Maida

Spiking neural networks (SNNs) enable power-efficient implementations due to their sparse, spike-based coding scheme. This paper develops a bio-inspired SNN that uses unsupervised learning to extract discriminative featu…

Unsupervised Learning of Syntactic Structure with Invertible Neural Projections

2018-08-28 · EMNLP 2018 10 · Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick

Unsupervised learning of syntactic structure is typically performed using generative models with discrete latent variables and multinomial parameters. In most cases, these models have not leveraged continuous word repres…

Constituency Grammar InductionDependency ParsingPOSUnsupervised Dependency Parsing

wav2vec: Unsupervised Pre-training for Speech Recognition

2019-04-11 · Steffen Schneider, Alexei Baevski, Ronan Collobert, Michael Auli

We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve…

Binary ClassificationGeneral ClassificationSpeech RecognitionUnsupervised Pre-training

Modeling speech recognition and synthesis simultaneously: Encoding and decoding lexical and sublexical semantic information into speech with no access to speech data

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Human speakers encode information into raw speech which is then decoded by the listeners. This complex relationship between encoding (production) and decoding (perception) is often modeled separately. Here, we test how d…

speech-recognitionSpeech RecognitionSpeech Synthesis

Unspeech: Unsupervised Speech Context Embeddings

2018-04-18 · Benjamin Milde, Chris Biemann

We introduce "Unspeech" embeddings, which are based on unsupervised learning of context feature representations for spoken language. The embeddings were trained on up to 9500 hours of crawled English speech data without …

Clustering