Phone-level pronunciation scoring
1개 벤치마크 · 논문 10편 · 이 태스크의 논문 보기 →
Benchmarks
speechocean762
Most implemented
speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment
Hierarchical Pronunciation Assessment with Multi-Aspect Attention
Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
A transfer learning based approach for pronunciation scoring
Papers
ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency scor…
Contrastive LearningPhone-level pronunciation scoringregressionUtterance-level pronounciation scoring+1Fine-Tuning Self-Supervised Learning Models for End-to-End Pronunciation Scoring
Automatic pronunciation assessment models are regularly used in language learning applications. Common methodologies for pronunciation assessment use feature-based approaches, such as the Goodness-of-Pronunciation (GOP) …
Feature EngineeringPhone-level pronunciation scoringPhoneme RecognitionSelf-Supervised Learning+3A Hierarchical Context-aware Modeling Approach for Multi-aspect and Multi-granular Pronunciation Assessment
Automatic Pronunciation Assessment (APA) plays a vital role in Computer-assisted Pronunciation Training (CAPT) when evaluating a second language (L2) learner's speaking proficiency. However, an apparent downside of most …
Automatic Speech RecognitionMulti-Task LearningPhone-level pronunciation scoringSentence+2Hierarchical Pronunciation Assessment with Multi-Aspect Attention
Automatic pronunciation assessment is a major component of a computer-assisted pronunciation training system. To provide in-depth feedback, scoring pronunciation at various levels of granularity such as phoneme, word, an…
Multi-Task LearningPhone-level pronunciation scoringUtterance-level pronounciation scoringWord-level pronunciation scoringSpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation
The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronuncia…
Data AugmentationMulti-Task LearningPhone-level pronunciation scoring0/1 Deep Neural Networks via Block Coordinate Descent
The step function is one of the simplest and most natural activation functions for deep neural networks (DNNs). As it counts 1 for positive variables and 0 for others, its intrinsic characteristics (e.g., discontinuity a…
10-shot image generation16k2D Object Detection+92