paper-with-me

Papers

A transfer learning based approach for pronunciation scoring

2021-11-01 · Marcelo Sancinetti, Jazmin Vidal, Cyntia Bonomi, Luciana Ferrer

Phone-level pronunciation scoring is a challenging task, with performance far from that of human annotators. Standard systems generate a score for each phone in a phrase using models trained for automatic speech recognition (ASR) with native data only. Better performance has been shown when using systems that are trained specifically for the task using non-native data. Yet, such systems face the challenge that datasets labelled for this task are scarce and usually small. In this paper, we present a transfer learning-based approach that leverages a model trained for ASR, adapting it for the task of pronunciation scoring. We analyze the effect of several design choices and compare the performance with a state-of-the-art goodness of pronunciation (GOP) system. Our final system is 20% better than the GOP system on EpaDB, a database for pronunciation scoring research, for a cost function that prioritizes low rates of unnecessary corrections.

📄 PDF Abstract BibTeX arXiv:2111.00976

Code (1)

marcelosancinetti/epa-gop-pykaldi 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phone-level pronunciation scoringspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Fine-Tuning Self-Supervised Learning Models for End-to-End Pronunciation Scoring

2023-09-19 · IEEE Access 2023 9 · Ahmed I. Zahran, Aly A. Fahmy, Khaled T. Wassif, Hanaa Bayomi

Automatic pronunciation assessment models are regularly used in language learning applications. Common methodologies for pronunciation assessment use feature-based approaches, such as the Goodness-of-Pronunciation (GOP) …

Feature EngineeringPhone-level pronunciation scoringPhoneme RecognitionSelf-Supervised Learning+3

Context-aware Goodness of Pronunciation for Computer-Assisted Pronunciation Training

2020-08-19

Mispronunciation detection is an essential component of the Computer-Assisted Pronunciation Training (CAPT) systems. State-of-the-art mispronunciation detection models use Deep Neural Networks (DNN) for acoustic modeling…

Sentence

Automatic Pronunciation Scoring And Mispronunciation Detection Using CMUSphinx

2012-12-01 · WS 2012 12 · Ronanki Srikanth, Bo Li, James Salsman
Speech Recognition

An End-to-End Mispronunciation Detection System for L2 English Speech Leveraging Novel Anti-Phone Modeling

2020-05-25 · Bi-Cheng Yan, Meng-Che Wu, Hsiao-Tsung Hung, Berlin Chen

Mispronunciation detection and diagnosis (MDD) is a core component of computer-assisted pronunciation training (CAPT). Most of the existing MDD approaches focus on dealing with categorical errors (viz. one canonical phon…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Leveraging phone-level linguistic-acoustic similarity for utterance-level pronunciation scoring

2023-02-21 · Wei Liu, Kaiqi Fu, Xiaohai Tian, Shuju Shi 외

Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding …