Learning robust speech representation with an articulatory-regularized variational autoencoder
It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingSpeech DenoisingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE
The human perception system is often assumed to recruit motor knowledge when processing auditory speech inputs. Using articulatory modeling and deep learning, this study examines how this articulatory information can be …
Learning to Compute the Articulatory Representations of Speech with the MIRRORNET
Most organisms including humans function by coordinating and integrating sensory signals with motor actions to survive and accomplish desired tasks. Learning these complex sensorimotor mappings proceeds simultaneously an…
Masked Autoencoders Are Articulatory Learners
Articulatory recordings track the positions and motion of different articulators along the vocal tract and are widely used to study speech production and to develop speech technologies such as articulatory based speech s…
Learning Joint Articulatory-Acoustic Representations with Normalizing Flows
The articulatory geometric configurations of the vocal tract and the acoustic properties of the resultant speech sound are considered to have a strong causal relationship. This paper aims at finding a joint latent repres…
Accent Conversion with Articulatory Representations
Conversion of non-native accented speech to native (American) English has a wide range of applications such as improving intelligibility of non-native speech. Previous work on this domain has used phonetic posteriograms …
Multi-Task Learning