Papers Phone-level pronunciation scoring
“Phone-level pronunciation scoring” 태그가 달린 논문 10편 · 필터 해제
ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency scor…
Contrastive LearningPhone-level pronunciation scoringregressionUtterance-level pronounciation scoring+1Fine-Tuning Self-Supervised Learning Models for End-to-End Pronunciation Scoring
Automatic pronunciation assessment models are regularly used in language learning applications. Common methodologies for pronunciation assessment use feature-based approaches, such as the Goodness-of-Pronunciation (GOP) …
Feature EngineeringPhone-level pronunciation scoringPhoneme RecognitionSelf-Supervised Learning+3A Hierarchical Context-aware Modeling Approach for Multi-aspect and Multi-granular Pronunciation Assessment
Automatic Pronunciation Assessment (APA) plays a vital role in Computer-assisted Pronunciation Training (CAPT) when evaluating a second language (L2) learner's speaking proficiency. However, an apparent downside of most …
Automatic Speech RecognitionMulti-Task LearningPhone-level pronunciation scoringSentence+2Hierarchical Pronunciation Assessment with Multi-Aspect Attention
Automatic pronunciation assessment is a major component of a computer-assisted pronunciation training system. To provide in-depth feedback, scoring pronunciation at various levels of granularity such as phoneme, word, an…
Multi-Task LearningPhone-level pronunciation scoringUtterance-level pronounciation scoringWord-level pronunciation scoringSpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation
The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronuncia…
Data AugmentationMulti-Task LearningPhone-level pronunciation scoring0/1 Deep Neural Networks via Block Coordinate Descent
The step function is one of the simplest and most natural activation functions for deep neural networks (DNNs). As it counts 1 for positive variables and 0 for others, its intrinsic characteristics (e.g., discontinuity a…
10-shot image generation16k2D Object Detection+92Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous eff…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task LearningPhone-level pronunciation scoring+4CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis
Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems. End-to-end (e2e) approaches are becoming dominant in MDD. However an e2e MDD model usual…
Multi-Task LearningPhone-level pronunciation scoringA transfer learning based approach for pronunciation scoring
Phone-level pronunciation scoring is a challenging task, with performance far from that of human annotators. Standard systems generate a score for each phone in a phrase using models trained for automatic speech recognit…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phone-level pronunciation scoringspeech-recognition+2speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment
This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are c…
Phone-level pronunciation scoringSentencespeech-recognition