A Deep Paradigm for Articulatory Speech Representation Learning via Neural Convolutive Sparse Matrix Factorization
Most of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure. This work, investigating the speech representations derived from articulatory kinematics signals, uses a neural implementation of convolutive sparse matrix factorization to decompose the articulatory data into interpretable gestures and gestural scores. By applying sparse constraints, the gestural scores leverage the discrete combinatorial properties of phonological gestures. Phoneme recognition experiments were additionally performed to show that gestural scores indeed code phonological information successfully. The proposed work thus makes a bridge between articulatory phonology and deep neural networks to leverage interpretable, intelligible, informative, and efficient speech representations.
Code (0)
등록된 구현이 없습니다.
Tasks
Phoneme RecognitionRepresentation LearningSpeech Representation LearningSimilar Papers 제목 키워드 기반
Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition
Most of the research on data-driven speech representation learning has focused on raw audios in an end-to-end manner, paying little attention to their internal phonological or gestural structure. This work, investigating…
Phoneme RecognitionRepresentation LearningSpeech Representation LearningArticulatory Representation Learning Via Joint Factor Analysis and Neural Matrix Factorization
Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures,…
Representation LearningArticulation GAN: Unsupervised modeling of articulatory learning
Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results i…
Generative Adversarial NetworkSpeech SynthesisSelf-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE
The human perception system is often assumed to recruit motor knowledge when processing auditory speech inputs. Using articulatory modeling and deep learning, this study examines how this articulatory information can be …
Learning to Compute the Articulatory Representations of Speech with the MIRRORNET
Most organisms including humans function by coordinating and integrating sensory signals with motor actions to survive and accomplish desired tasks. Learning these complex sensorimotor mappings proceeds simultaneously an…