paper-with-me

홈 › Papers

Automated detection of pronunciation errors in non-native English speech employing deep learning

2022-09-13 · Daniel Korzekwa

Despite significant advances in recent years, the existing Computer-Assisted Pronunciation Training (CAPT) methods detect pronunciation errors with a relatively low accuracy (precision of 60% at 40%-80% recall). This Ph.D. work proposes novel deep learning methods for detecting pronunciation errors in non-native (L2) English speech, outperforming the state-of-the-art method in AUC metric (Area under the Curve) by 41%, i.e., from 0.528 to 0.749. One of the problems with existing CAPT methods is the low availability of annotated mispronounced speech needed for reliable training of pronunciation error detection models. Therefore, the detection of pronunciation errors is reformulated to the task of generating synthetic mispronounced speech. Intuitively, if we could mimic mispronounced speech and produce any amount of training data, detecting pronunciation errors would be more effective. Furthermore, to eliminate the need to align canonical and recognized phonemes, a novel end-to-end multi-task technique to directly detect pronunciation errors was proposed. The pronunciation error detection models have been used at Amazon to automatically detect pronunciation errors in synthetic speech to accelerate the research into new speech synthesis methods. It was demonstrated that the proposed deep learning methods are applicable in the tasks of detecting and reconstructing dysarthric speech.

📄 PDF Abstract BibTeX arXiv:2209.06265

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Experiments of ASR-based mispronunciation detection for children and adult English learners

2021-04-13 · Nina Hosseini-Kivanani, Roberto Gretter, Marco Matassoni, Giuseppe Daniele Falavigna

Pronunciation is one of the fundamentals of language learning, and it is considered a primary factor of spoken language when it comes to an understanding and being understood by others. The persistent presence of high er…

Language Modellingspeech-recognitionSpeech Recognition

AUTOMATIC PRONUNCIATION MISTAKE DETECTOR PROJECT REPORT

2025-06-25 · ResearchGate 2025 6 · Kamal Acharya

Given the drawbacks of traditional English pronunciation correction systems, such as failure to provide timely feedback and correct learners' pronunciation errors, slow improvement of learners' English proficiency, and e…

Mistake Detectionspeech-recognitionSpeech Recognition

Weakly-supervised word-level pronunciation error detection in non-native English speech

2021-06-07 · Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman, Shira Calamaro 외

We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not required and we only need to mark mispronou…

Computer-assisted Pronunciation Training -- Speech synthesis is almost all you need

2022-07-02 · Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman, Bozena Kostek

The research community has long studied computer-assisted pronunciation training (CAPT) methods in non-native speech. Researchers focused on studying various model architectures, such as Bayesian networks and deep learni…

AllSpeech Synthesistext-to-speechText to Speech

Phonological Level wav2vec2-based Mispronunciation Detection and Diagnosis Method

2023-11-13 · Mostafa Shahin, Julien Epps, Beena Ahmed

The automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Languag…

AttributeDiagnostic