MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Task LearningSentenceSimilar Papers 제목 키워드 기반
Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous eff…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task LearningPhone-level pronunciation scoring+4Multi-task Pretraining for Enhancing Interpretable L2 Pronunciation Assessment
Automatic pronunciation assessment (APA) analyzes second-language (L2) learners' speech by providing fine-grained pronunciation feedback at various linguistic levels. Most existing efforts on APA typically adopt segmenta…
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the ca…
Fine-Tuning Self-Supervised Learning Models for End-to-End Pronunciation Scoring
Automatic pronunciation assessment models are regularly used in language learning applications. Common methodologies for pronunciation assessment use feature-based approaches, such as the Goodness-of-Pronunciation (GOP) …
Feature EngineeringPhone-level pronunciation scoringPhoneme RecognitionSelf-Supervised Learning+3Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
In automated pronunciation assessment, recent emphasis progressively lies on evaluating multiple aspects to provide enriched feedback. However, acquiring multi-aspect-score labeled data for non-native language learners' …
speech-recognitionSpeech Recognition