The Use of Voice Source Features for Sung Speech Recognition
In this paper, we ask whether vocal source features (pitch, shimmer, jitter, etc) can improve the performance of automatic sung speech recognition, arguing that conclusions previously drawn from spoken speech studies may not be valid in the sung speech domain. We first use a parallel singing/speaking corpus (NUS-48E) to illustrate differences in sung vs spoken voicing characteristics including pitch range, syllables duration, vibrato, jitter and shimmer. We then use this analysis to inform speech recognition experiments on the sung speech DSing corpus, using a state of the art acoustic model and augmenting conventional features with various voice source parameters. Experiments are run with three standard (increasingly large) training sets, DSing1 (15.1 hours), DSing3 (44.7 hours) and DSing30 (149.1 hours). Pitch combined with degree of voicing produces a significant decrease in WER from 38.1% to 36.7% when training with DSing1 however smaller decreases in WER observed when training with the larger more varied DSing3 and DSing30 sets were not seen to be statistically significant. Voicing quality characteristics did not improve recognition performance although analysis suggests that they do contribute to an improved discrimination between voiced/unvoiced phoneme pairs.
Code (0)
등록된 구현이 없습니다.
Tasks
speech-recognitionSpeech RecognitionvalidSimilar Papers 제목 키워드 기반
Convolutional Speech Recognition with Pitch and Voice Quality Features
The effects of adding pitch and voice quality features such as jitter and shimmer to a state-of-the-art CNN model for Automatic Speech Recognition are studied in this work. Pitch features have been previously used for im…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionspeech-recognition+1A Singing Voice Database in Basque for Statistical Singing Synthesis of Bertsolaritza
This paper describes the characteristics and structure of a Basque singing voice database of bertsolaritza. Bertsolaritza is a popular singing style from Basque Country sung exclusively in Basque that is improvised and a…
Singing Voice SynthesisSpeaker Recognition in Bengali Language from Nonlinear Features
At present Automatic Speaker Recognition system is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking style of a pers…
Speaker IdentificationSpeaker Recognitionspeech-recognitionSpeech RecognitionMacsen: A Voice Assistant for Speakers of a Lesser Resourced Language
This paper reports on the development of a voice assistant mobile app for speakers of a lesser resourced language {--} Welsh. An assistant with a smaller set of effective but useful skills is both desirable and urgent fo…
Language Modelingspeech-recognitionSpeech RecognitionTransfer LearningBengali Common Voice Speech Dataset for Automatic Speech Recognition
Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversitySentence+2