Reading Miscue Detection in Primary School through Automatic Speech Recognition
Automatic reading diagnosis systems can benefit both teachers for more efficient scoring of reading exercises and students for accessing reading exercises with feedback more easily. However, there are limited studies on Automatic Speech Recognition (ASR) for child speech in languages other than English, and limited research on ASR-based reading diagnosis systems. This study investigates how efficiently state-of-the-art (SOTA) pretrained ASR models recognize Dutch native children speech and manage to detect reading miscues. We found that Hubert Large finetuned on Dutch speech achieves SOTA phoneme-level child speech recognition (PER at 23.1\%), while Whisper (Faster Whisper Large-v2) achieves SOTA word-level performance (WER at 9.8\%). Our findings suggest that Wav2Vec2 Large and Whisper are the two best ASR models for reading miscue detection. Specifically, Wav2Vec2 Large shows the highest recall at 0.83, whereas Whisper exhibits the highest precision at 0.52 and an F1 score of 0.52.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection
Identifying mistakes (i.e., miscues) made while reading aloud is commonly approached post-hoc by comparing automatic speech recognition (ASR) transcriptions to the target reading text. However, post-hoc methods perform p…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionThe Effect of Teacher Gender on Student Achievement in Primary School
Using data from a randomized experiment, we find that having a fe- male teacher lowers the math test scores of female primary school students in disadvantaged neighborhoods. Moreover, we do not find any effect of having …
MathThe LetsRead Corpus of Portuguese Children Reading Aloud for Performance Evaluation
This paper introduces the LetsRead Corpus of European Portuguese read speech from 6 to 10 years old children. The motivation for the creation of this corpus stems from the inexistence of databases with recordings of read…
Towards Multi-Modal Text-Image Retrieval to improve Human Reading
In primary school, children{'}s books, as well as in modern language learning apps, multi-modal learning strategies like illustrations of terms and phrases are used to support reading comprehension. Also, several studies…
Image RetrievalReading ComprehensionRetrievalChanging views about remote working during the COVID-19 pandemic: Evidence using panel data from Japan
COVID-19 has led to school closures in Japan to cope with the pandemic. Under the state of emergency, in addition to school closure, after-school care has not been sufficiently supplied. We independently collected indivi…