paper-with-me

Papers

Error-preserving Automatic Speech Recognition of Young English Learners' Language

2024-06-05 · Janick Michot, Manuela Hürlimann, Jan Deriu, Luzia Sauer, Katsiaryna Mlynchyk, Mark Cieliebak

One of the central skills that language learners need to practice is speaking the language. Currently, students in school do not get enough speaking opportunities and lack conversational practice. Recent advances in speech technology and natural language processing allow for the creation of novel tools to practice their speaking skills. In this work, we tackle the first component of such a pipeline, namely, the automated speech recognition module (ASR), which faces a number of challenges: first, state-of-the-art ASR models are often trained on adult read-aloud data by native speakers and do not transfer well to young language learners' speech. Second, most ASR systems contain a powerful language model, which smooths out errors made by the speakers. To give corrective feedback, which is a crucial part of language learning, the ASR systems in our setting need to preserve the errors made by the language learners. In this work, we build an ASR system that satisfies these requirements: it works on spontaneous speech by young language learners and preserves their errors. For this, we collected a corpus containing around 85 hours of English audio spoken by learners in Switzerland from grades 4 to 6 on different language learning tasks, which we used to train an ASR model. Our experiments show that our model benefits from direct fine-tuning on children's voices and has a much higher error preservation rate than other models.

📄 PDF Abstract BibTeX arXiv:2406.03235

Code (1)

mict-zhaw/chall_e2e_stt 공식 구현 pytorch

Tasks

Automatic Speech RecognitionLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Fundamental Frequency Feature Normalization and Data Augmentation for Child Speech Recognition

2021-02-18 · Gary Yeung, Ruchao Fan, Abeer Alwan

Automatic speech recognition (ASR) systems for young children are needed due to the importance of age-appropriate educational technology. Because of the lack of publicly available young child speech data, feature extract…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1

End-to-end acoustic modelling for phone recognition of young readers

2021-03-04 · Lucile Gelin, Morgane Daniel, Julien Pinquier, Thomas Pellegrini

Automatic recognition systems for child speech are lagging behind those dedicated to adult speech in the race of performance. This phenomenon is due to the high acoustic and linguistic variability present in child speech…

Acoustic ModellingTransfer Learning

You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish

2024-05-22 · Ronald Cumbal, Birger Moell, Jose Lopes, Olof Engwall

The performance of Automatic Speech Recognition (ASR) systems has constantly increased in state-of-the-art development. However, performance tends to decrease considerably in more challenging conditions (e.g., background…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale

2023-04-30 · Hiuchung Hung, Andreas Maier, Thorsten Piske

This paper introduces a non-native speech corpus consisting of narratives from fifty 5- to 6-year-old Chinese-English children. Transcripts totaling 6.5 hours of children taking a narrative comprehension test in English …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

CORAA: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese

2021-10-14 · Arnaldo Candido Junior, Edresson Casanova, Anderson Soares, Frederico Santos de Oliveira 외

Automatic Speech recognition (ASR) is a complex and challenging task. In recent years, there have been significant advances in the area. In particular, for the Brazilian Portuguese (BP) language, there were about 376 hou…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition