Speaker conditioned acoustic-to-articulatory inversion using x-vectors
Speech production involves the movement of various articulators, including tongue, jaw, and lips. Estimating the movement of the articulators from the acoustics of speech is known as acoustic-to-articulatory inversion (AAI). Recently, it has been shown that instead of training AAI in a speaker specific manner, pooling the acoustic-articulatory data from multiple speakers is beneficial. Further, additional conditioning with speaker specific information by one-hot encoding at the input of AAI along with acoustic features benefits the AAI performance in a closed-set speaker train and test condition. In this work, we carry out an experimental study on the benefit of using x-vectors for providing speaker specific information to condition AAI. Experiments with 30 speakers have shown that the AAI performance benefits from the use of x-vectors in a closed set seen speaker condition. Further, x-vectors also generalizes well for unseen speaker evaluation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Articulatory-WaveNet: Autoregressive Model For Acoustic-to-Articulatory Inversion
This paper presents Articulatory-WaveNet, a new approach for acoustic-to-articulator inversion. The proposed system uses the WaveNet speech synthesis architecture, with dilated causal convolutional layers using previous …
Speech SynthesisVers une inversion acoustico-articulatoire d'un locuteur \'etranger (Toward an acoustic to articulatory inversion of a foreign speaker) [in French]
Speaker-Independent Acoustic-to-Articulatory Speech Inversion
To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is…
ResynthesisExploiting Cross Domain Acoustic-to-articulatory Inverted Features For Disordered Speech Recognition
Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disor…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1Training Articulatory Inversion Models for Interspeaker Consistency
Acoustic-to-Articulatory Inversion (AAI) attempts to model the inverse mapping from speech to articulation. Exact articulatory prediction from speech alone may be impossible, as speakers can choose different forms of art…
Self-Supervised Learning