LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
Prevalent ungrammatical expressions and disfluencies in spontaneous speech from second language (L2) learners pose unique challenges to Automatic Speech Recognition (ASR) systems. However, few datasets are tailored to L2 learner speech. We publicly release LearnerVoice, a dataset consisting of 50.04 hours of audio and transcriptions of L2 learners' spontaneous speech. Our linguistic analysis reveals that transcriptions in our dataset contain L2S (L2 learner's Spontaneous speech) features, consisting of ungrammatical expressions and disfluencies (e.g., filler words, word repetitions, self-repairs, false starts), significantly more than native speech datasets. Fine-tuning whisper-small.en with LearnerVoice achieves a WER of 10.26%, 44.2% lower than vanilla whisper-small.en. Furthermore, our qualitative analysis indicates that 54.2% of errors from the vanilla model on LearnerVoice are attributable to L2S features, with 48.1% of them being reduced in the fine-tuned model.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Error-preserving Automatic Speech Recognition of Young English Learners' Language
One of the central skills that language learners need to practice is speaking the language. Currently, students in school do not get enough speaking opportunities and lack conversational practice. Recent advances in spee…
Automatic Speech RecognitionLanguage Modellingspeech-recognitionSpeech RecognitionUsing Rhetorical Structure Theory to Assess Discourse Coherence for Non-native Spontaneous Speech
This study aims to model the discourse structure of spontaneous spoken responses within the context of an assessment of English speaking proficiency for non-native speakers. Rhetorical Structure Theory (RST) has been com…
To What Extent Does Lexical Normalization Help English-as-a-Second Language Learners to Read Noisy English Texts?
How difficult is it for English-as-a-second language (ESL) learners to read noisy English texts? Do ESL learners need lexical normalization to read noisy English texts? These questions may also affect community formation…
Lexical NormalizationToward Automated Content Feedback Generation for Non-native Spontaneous Speech
In this study, we developed an automated algorithm to provide feedback about the specific content of non-native English speakers{'} spoken responses. The responses were spontaneous speech, elicited using integrated tasks…
speech-recognitionSpeech RecognitionA Preliminary Study on Automated Speaking Assessment of English as a Second Language (ESL) Students
Due to the surge in global demand for English as a second language (ESL), developments of automated methods for grading speaking proficiency have gained considerable attention. This paper aims to present a computerized r…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition