paper-with-me

Papers

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

2026-06-23 · Milosz Dudek, Daria Hemmerling, Kamil Kwarciak, Maciej Stroinski, Maria Pensko, Mateusz Kowalewski, Leonid Pavlovskyi, Sebastian Jurczak, Anna-Mariia Vitkovska, Zuzanna Miodonska, Natalia Mocko, Michal Krecichwost arxiv

Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can operate outside the clinic. We present a screening pipeline for Polish-speaking children focused on sibilant substitutions, coupling a wav2vec2-based CTC token recognizer with alignment-based error typing and a template-grounded caregiver assistant for screening, not diagnosis. On a held-out test set of 10 unseen children comprising 559 utterances, the recognizer achieves 88.7 percent exact sequence match. As a conservative screening proxy, we flag a mismatch when the system emits substitution-evidence bracketed tokens at the target segment, yielding 72.9 percent precision, 61.4 percent recall, F1 = 0.67, and a 2.7 percent false-alarm rate on target-correct items. We describe the assistant's safety boundaries and outline a clinician-in-the-loop validation plan for future deployment.

📄 PDF Abstract BibTeX arXiv:2606.25181

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mispronunciation Detection in Non-native (L2) English with Uncertainty Modeling

2021-01-16 · Daniel Korzekwa, Jaime Lorenzo-Trueba, Szymon Zaporowski, Shira Calamaro 외

A common approach to the automatic detection of mispronunciation in language learning is to recognize the phonemes produced by a student and compare it to the expected pronunciation of a native speaker. This approach mak…

Automatic Phoneme RecognitionPhoneme RecognitionSentencevalid

Weakly-supervised word-level pronunciation error detection in non-native English speech

2021-06-07 · Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman, Shira Calamaro 외

We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not required and we only need to mark mispronou…

SpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation

2022-11-02 · Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali, Hamdy Mubarak 외

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronuncia…

Data AugmentationMulti-Task LearningPhone-level pronunciation scoring

Phonological Level wav2vec2-based Mispronunciation Detection and Diagnosis Method

2023-11-13 · Mostafa Shahin, Julien Epps, Beena Ahmed

The automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Languag…

AttributeDiagnostic

Context-aware Goodness of Pronunciation for Computer-Assisted Pronunciation Training

2020-08-19

Mispronunciation detection is an essential component of the Computer-Assisted Pronunciation Training (CAPT) systems. State-of-the-art mispronunciation detection models use Deep Neural Networks (DNN) for acoustic modeling…

Sentence