paper-with-me

Papers

Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition

2025-02-01 · Anna Seo Gyeong Choi, JongHyeon Park, Myungwoo Oh

Recent advancements in machine learning have significantly improved speech recognition, but recognizing speech from non-fluent or accented speakers remains a challenge. Previous efforts, relying on rule-based pronunciation patterns, have struggled to fully capture non-native errors. We propose two data-driven approaches using speech corpora to automatically detect mispronunciation patterns. By aligning non-native phones with their native counterparts using attention maps, we achieved a 5.7% improvement in speech recognition on native English datasets and a 12.8% improvement for non-native English speakers, particularly Korean speakers. Our method offers practical advancements for robust Automatic Speech Recognition (ASR) systems particularly for situations where prior linguistic knowledge is not applicable.

📄 PDF Abstract BibTeX arXiv:2502.00583

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Acoustic-to-Articulatory Speech Inversion Features for Mispronunciation Detection of /r/ in Child Speech Sound Disorders

2023-05-25 · Nina R Benway, Yashish M Siriwardena, Jonathan L Preston, Elaine Hitchcock 외

Acoustic-to-articulatory speech inversion could enhance automated clinical mispronunciation detection to provide detailed articulatory feedback unattainable by formant-based mispronunciation detection algorithms; however…

Towards Temporally Explainable Dysarthric Speech Clarity Assessment

2025-05-31 · Seohyun Park, Chitralekha Gupta, Michelle Kah Yian Kwan, Xinhui Fung 외

Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

SpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation

2022-11-02 · Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali, Hamdy Mubarak 외

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronuncia…

Data AugmentationMulti-Task LearningPhone-level pronunciation scoring

Exploring the Integration of E2E ASR and Pronunciation Modeling for English Mispronunciation Detection

2021-10-01 · ROCLING 2021 10 · Hsin-Wei Wang, Bi-Cheng Yan, Yung-Chang Hsu, Berlin Chen

There has been increasing demand to develop effective computer-assisted language training (CAPT) systems, which can provide feedback on mispronunciations and facilitate second-language (L2) learners to improve their spea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Deep segmental phonetic posterior-grams based discovery of non-categories in L2 English speech

2020-02-01 · Xu Li, Xixin Wu, Xunying Liu, Helen Meng

Second language (L2) speech is often labeled with the native, phone categories. However, in many cases, it is difficult to decide on a categorical phone that an L2 segment belongs to. These segments are regarded as non-c…