paper-with-me

Papers

A comparative study of estimating articulatory movements from phoneme sequences and acoustic features

2019-10-31 · Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh

Unlike phoneme sequences, movements of speech articulators (lips, tongue, jaw, velum) and the resultant acoustic signal are known to encode not only the linguistic message but also carry para-linguistic information. While several works exist for estimating articulatory movement from acoustic signals, little is known to what extent articulatory movements can be predicted only from linguistic information, i.e., phoneme sequence. In this work, we estimate articulatory movements from three different input representations: R1) acoustic signal, R2) phoneme sequence, R3) phoneme sequence with timing information. While an attention network is used for estimating articulatory movement in the case of R2, BLSTM network is used for R1 and R3. Experiments with ten subjects' acoustic-articulatory data reveal that the estimation techniques achieve an average correlation coefficient of 0.85, 0.81, and 0.81 in the case of R1, R2, and R3 respectively. This indicates that attention network, although uses only phoneme sequence (R2) without any timing information, results in an estimation performance similar to that using rich acoustic signal (R1), suggesting that articulatory motion is primarily driven by the linguistic message. The correlation coefficient is further improved to 0.88 when R1 and R3 are used together for estimating articulatory movements.

📄 PDF Abstract BibTeX arXiv:1910.14375

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech

2024-07-03 · Tobias Weise, Philipp Klumpp, Kubilay Can Demir, Paula Andrea Pérez-Toro 외

This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as a…

Motion EstimationMulti-Task Learning

Estimating articulatory movements in speech production with transformer networks

2021-04-11 · Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa 외

We estimate articulatory movements in speech production from different modalities - acoustics and phonemes. Acoustic-to articulatory inversion (AAI) is a sequence-to-sequence task. On the other hand, phoneme to articulat…

Motion Estimation

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

2026-04-20 · Chenqian Le, Ruisi Li, Beatrice Fumagalli, Yasamin Esmaeili 외 arxiv

We test whether Speech Articulatory Coding (SPARC) features can linearly predict surface electromyography (sEMG) envelopes across aloud, mimed, and subvocal speech in twenty-four subjects. Using elastic-net multivariate …

Attention and Encoder-Decoder based models for transforming articulatory movements at different speaking rates

2020-06-04 · Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh

While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform ar…

Decoder

Latent semantics of action verbs reflect phonetic parameters of intensity and emotional content

2014-05-06 · Michael Kai Petersen

Conjuring up our thoughts, language reflects statistical patterns of word co-occurrences which in turn come to describe how we perceive the world. Whether counting how frequently nouns and verbs combine in Google search …

Clustering