paper-with-me

Papers

Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images

2020-06-29 · Pramit Saha, Yadong Liu, Bryan Gick, Sidney Fels

Thousands of individuals need surgical removal of their larynx due to critical diseases every year and therefore, require an alternative form of communication to articulate speech sounds after the loss of their voice box. This work addresses the articulatory-to-acoustic mapping problem based on ultrasound (US) tongue images for the development of a silent-speech interface (SSI) that can provide them with an assistance in their daily interactions. Our approach targets automatically extracting tongue movement information by selecting an optimal feature set from US images and mapping these features to the acoustic space. We use a novel deep learning architecture to map US tongue images from the US probe placed beneath a subject's chin to formants that we call, Ultrasound2Formant (U2F) Net. It uses hybrid spatio-temporal 3D convolutions followed by feature shuffling, for the estimation and tracking of vowel formants from US images. The formant values are then utilized to synthesize continuous time-varying vowel trajectories, via Klatt Synthesizer. Our best model achieves R-squared (R^2) measure of 99.96% for the regression task. Our network lays the foundation for an SSI as it successfully tracks the tongue contour automatically as an internal representation without any explicit annotation.

📄 PDF Abstract BibTeX arXiv:2006.16367

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Use of a Spectral Glottal Model for the Source-filter Separation of Speech

2017-12-21

The estimation of glottal flow from a speech waveform is a key method for speech analysis and parameterization. Significant research effort has been made to dissociate the first vocal tract resonance from the glottal for…

Quantifying and Correlating Rhythm Formants in Speech

2019-09-03 · Dafydd Gibbon, Peng Li

The objective of the present study is exploratory: to introduce and apply a new theory of speech rhythm zones or rhythm formants (R-formants). R-formants are zones of high magnitude frequencies in the low frequency (LF) …

Rhythm

Distinguishable Speaker Anonymization based on Formant and Fundamental Frequency Scaling

2022-11-06 · Jixun Yao, Qing Wang, Yi Lei, Pengcheng Guo 외

Speech data on the Internet are proliferating exponentially because of the emergence of social media, and the sharing of such personal data raises obvious security and privacy concerns. One solution to mitigate these con…

Speaker anonymizationSpeaker Verification

A Wavelet Transform Based Scheme to Extract Speech Pitch and Formant Frequencies

2022-09-01 · Seyedamiryousef Hosseini Goki, Mahdieh Ghazvini, Sajad Hamzenejadi

Pitch and Formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are ess…

speech-recognitionSpeech Recognition

Exploring rhythm formant analysis for Indic language classification

2024-10-08 · Parismita Gogoi, Sishir Kalita, Priyankoo Sarmah, S. R Mahadeva Prasanna

This paper reports a preliminary study on quantitative frequency domain rhythm cues for classifying five Indian languages: Bengali, Kannada, Malayalam, Marathi, and Tamil. We employ rhythm formant (R-formants) analysis, …

ClassificationRhythm