paper-with-me

Papers

The Use of Voice Source Features for Sung Speech Recognition

2021-02-20 · Gerardo Roa Dabike, Jon Barker

In this paper, we ask whether vocal source features (pitch, shimmer, jitter, etc) can improve the performance of automatic sung speech recognition, arguing that conclusions previously drawn from spoken speech studies may not be valid in the sung speech domain. We first use a parallel singing/speaking corpus (NUS-48E) to illustrate differences in sung vs spoken voicing characteristics including pitch range, syllables duration, vibrato, jitter and shimmer. We then use this analysis to inform speech recognition experiments on the sung speech DSing corpus, using a state of the art acoustic model and augmenting conventional features with various voice source parameters. Experiments are run with three standard (increasingly large) training sets, DSing1 (15.1 hours), DSing3 (44.7 hours) and DSing30 (149.1 hours). Pitch combined with degree of voicing produces a significant decrease in WER from 38.1% to 36.7% when training with DSing1 however smaller decreases in WER observed when training with the larger more varied DSing3 and DSing30 sets were not seen to be statistically significant. Voicing quality characteristics did not improve recognition performance although analysis suggests that they do contribute to an improved discrimination between voiced/unvoiced phoneme pairs.

📄 PDF Abstract BibTeX arXiv:2102.10376

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognitionvalid

Similar Papers 제목 키워드 기반

Convolutional Speech Recognition with Pitch and Voice Quality Features

2020-09-02 · Guillermo Cámbara, Jordi Luque, Mireia Farrús

The effects of adding pitch and voice quality features such as jitter and shimmer to a state-of-the-art CNN model for Automatic Speech Recognition are studied in this work. Pitch features have been previously used for im…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionspeech-recognition+1

A Singing Voice Database in Basque for Statistical Singing Synthesis of Bertsolaritza

2016-05-01 · LREC 2016 5 · Xabier Sarasola, Eva Navas, David Tavarez, Daniel Erro 외

This paper describes the characteristics and structure of a Basque singing voice database of bertsolaritza. Bertsolaritza is a popular singing style from Basque Country sung exclusively in Basque that is improvised and a…

Singing Voice Synthesis

Speaker Recognition in Bengali Language from Nonlinear Features

2020-04-15 · Uddalok Sarkar, Soumyadeep Pal, Sayan Nag, Chirayata Bhattacharya 외

At present Automatic Speaker Recognition system is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking style of a pers…

Speaker IdentificationSpeaker Recognitionspeech-recognitionSpeech Recognition

Macsen: A Voice Assistant for Speakers of a Lesser Resourced Language

2020-05-01 · LREC 2020 5 · Dewi Jones

This paper reports on the development of a voice assistant mobile app for speakers of a lesser resourced language {--} Welsh. An assistant with a smaller set of effective but useful skills is both desirable and urgent fo…

Language Modelingspeech-recognitionSpeech RecognitionTransfer Learning

Bengali Common Voice Speech Dataset for Automatic Speech Recognition

2022-06-28 · Samiul Alam, Asif Sushmit, Zaowad Abdullah, Shahrin Nakkhatra 외

Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversitySentence+2