paper-with-me

Papers

Phone Duration Modeling for Speaker Age Estimation in Children

2021-09-03 · Prashanth Gurunath Shivakumar, Somer Bishop, Catherine Lord, Shrikanth Narayanan

Automatic inference of important paralinguistic information such as age from speech is an important area of research with numerous spoken language technology based applications. Speaker age estimation has applications in enabling personalization and age-appropriate curation of information and content. However, research in speaker age estimation in children is especially challenging due to paucity of relevant speech data representing the developmental spectrum, and the high signal variability especially intra age variability that complicates modeling. Most approaches in children speaker age estimation adopt methods directly from research on adult speech processing. In this paper, we propose features specific to children and focus on speaker's phone duration as an important biomarker of children's age. We propose phone duration modeling for predicting age from child's speech. To enable that, children speech is first forced aligned with the corresponding transcription to derive phone duration distributions. Statistical functionals are computed from phone duration distributions for each phoneme which are in turn used to train regression models to predict speaker age. Two children speech datasets are employed to demonstrate the robustness of phone duration features. We perform age regression experiments on age categories ranging from children studying in kindergarten to grade 10. Experimental results suggest phone durations contain important development-related information of children. Phonemes contributing most to estimation of children speaker age are analyzed and presented.

📄 PDF Abstract BibTeX arXiv:2109.01568

Code (0)

등록된 구현이 없습니다.

Tasks

Age Estimationregression

Similar Papers 제목 키워드 기반

Neural Network-Based Modeling of Phonetic Durations

2019-09-06 · Xizi Wei, Melvyn Hunt, Adrian Skilling

A deep neural network (DNN)-based model has been developed to predict non-parametric distributions of durations of phonemes in specified phonetic contexts and used to explore which factors influence durations most. Major…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Total-Duration-Aware Duration Modeling for Text-to-Speech Systems

2024-06-06 · Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Chung-Hsien Tsai 외

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such a…

Diversitytext-to-speechText to Speech

Filter-based Discriminative Autoencoders for Children Speech Recognition

2022-04-01 · Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao, Hsin-Min Wang

Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acoustic modeling. To filter out the influen…

DecoderDiversityDomain Adaptationspeech-recognition+1

Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization

2024-12-22 · Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi

In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verific…

Speaker Verification

Speaker- and Age-Invariant Training for Child Acoustic Modeling Using Adversarial Multi-Task Learning

2022-10-19 · Mostafa Shahin, Beena Ahmed, Julien Epps

One of the major challenges in acoustic modelling of child speech is the rapid changes that occur in the children's articulators as they grow up, their differing growth rates and the subsequent high variability in the sa…

Acoustic ModellingMulti-Task Learningspeech-recognitionSpeech Recognition