paper-with-me

Papers

speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment

2021-04-03 · Junbo Zhang, Zhiwen Zhang, Yongqing Wang, Zhiyong Yan, Qiong Song, YuKai Huang, Ke Li, Daniel Povey, Yujun Wang

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children. Five experts annotated each of the utterances at sentence-level, word-level and phoneme-level. A baseline system is released in open source to illustrate the phoneme-level pronunciation assessment workflow on this corpus. This corpus is allowed to be used freely for commercial and non-commercial purposes. It is available for free download from OpenSLR, and the corresponding baseline system is published in the Kaldi speech recognition toolkit.

📄 PDF Abstract BibTeX arXiv:2104.01378

Code (2)

kaldi-asr/kaldi/tree/master/egs/gop_speechocean762 공식 구현
YuanGongND/gopt pytorch

Tasks

Phone-level pronunciation scoringSentencespeech-recognition

Similar Papers 제목 키워드 기반

Goodness-of-pronunciation without phoneme time alignment

2026-03-26 · Jeremy H. M. Wong, Nancy F. Chen arxiv

In speech evaluation, an Automatic Speech Recognition (ASR) model often computes time boundaries and phoneme posteriors for input features. However, limited data for ASR training hinders expansion of speech evaluation to…

Speech Recognition

Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment

2022-05-06 · Yuan Gong, Ziyi Chen, Iek-Heng Chu, Peng Chang 외

Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous eff…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task LearningPhone-level pronunciation scoring+4

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

2026-06-18 · Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury arxiv

Training automated pronunciation assessment often relies on labeled learner errors or non-native corpora that are costly to collect. We propose a lightweight framework trained only on native speech resources, operating u…

L1-aware Multilingual Mispronunciation Detection Framework

2023-09-14 · Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali

The phonological discrepancies between a speaker's native (L1) and the non-native language (L2) serves as a major factor for mispronunciation. This paper introduces a novel multilingual MDD architecture, L1-MultiMDD, enr…

Phoneme Recognition

Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment

2025-03-14 · Ke Wang, Lei He, Kun Liu, Yan Deng 외

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the ca…