paper-with-me

Papers

SpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation

2022-11-02 · Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali, Hamdy Mubarak, Shazia Afzal

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronunciation errors to overcome such data scarcity. The SpeechBlender utilizes varieties of masks to target different regions of phonetic units, and use the mixing factors to linearly interpolate raw speech signals while augmenting pronunciation. The masks facilitate smooth blending of the signals, generating more effective samples than the `Cut/Paste' method. Our proposed technique achieves state-of-the-art results, with Speechocean762, on ASR dependent mispronunciation detection models at phoneme level, with a 2.0% gain in Pearson Correlation Coefficient (PCC) compared to the previous state-of-the-art [1]. Additionally, we demonstrate a 5.0% improvement at the phoneme level compared to our baseline. We also observed a 4.6% increase in F1-score with Arabic AraVoiceL2 testset.

📄 PDF Abstract BibTeX arXiv:2211.00923

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationMulti-Task LearningPhone-level pronunciation scoring

Similar Papers 제목 키워드 기반

Towards Temporally Explainable Dysarthric Speech Clarity Assessment

2025-05-31 · Seohyun Park, Chitralekha Gupta, Michelle Kah Yian Kwan, Xinhui Fung 외

Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Acoustic-to-Articulatory Speech Inversion Features for Mispronunciation Detection of /r/ in Child Speech Sound Disorders

2023-05-25 · Nina R Benway, Yashish M Siriwardena, Jonathan L Preston, Elaine Hitchcock 외

Acoustic-to-articulatory speech inversion could enhance automated clinical mispronunciation detection to provide detailed articulatory feedback unattainable by formant-based mispronunciation detection algorithms; however…

An Effective End-to-End Modeling Approach for Mispronunciation Detection

2020-05-18 · Tien-Hong Lo, Shi-Yan Weng, Hsiu-jui Chang, Berlin Chen

Recently, end-to-end (E2E) automatic speech recognition (ASR) systems have garnered tremendous attention because of their great success and unified modeling paradigms in comparison to conventional hybrid DNN-HMM ASR syst…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decision Makingspeech-recognition+1

AraS2P: Arabic Speech-to-Phonemes System

2025-09-27 · Bassam Matar, Mohamed Fayed, Ayman Khalafallah arxiv

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was…

CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis

2021-11-16 · Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung 외

Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems. End-to-end (e2e) approaches are becoming dominant in MDD. However an e2e MDD model usual…

Multi-Task LearningPhone-level pronunciation scoring