paper-with-me

Papers

Computational Pronunciation Analysis in Sung Utterances

2021-06-21 · Emir Demirel, Sven Ahlback, Simon Dixon

Recent automatic lyrics transcription (ALT) approaches focus on building stronger acoustic models or in-domain language models, while the pronunciation aspect is seldom touched upon. This paper applies a novel computational analysis on the pronunciation variances in sung utterances and further proposes a new pronunciation model adapted for singing. The singing-adapted model is tested on multiple public datasets via word recognition experiments. It performs better than the standard speech dictionary in all settings reporting the best results on ALT in a capella recordings using n-gram language models. For reproducibility, we share the sentence-level annotations used in testing, providing a new benchmark evaluation set for ALT.

📄 PDF Abstract BibTeX arXiv:2106.10977

Code (1)

emirdemirel/ALTA 공식 구현

Tasks

Automatic Lyrics TranscriptionSentence

Similar Papers 제목 키워드 기반

Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison

2025-07-15 · Andrew Valdivia, Yueming Zhang, Hailu Xu, Amir Ghasemkhani 외

This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions w…

Voice Cloning

Experiments of ASR-based mispronunciation detection for children and adult English learners

2021-04-13 · Nina Hosseini-Kivanani, Roberto Gretter, Marco Matassoni, Giuseppe Daniele Falavigna

Pronunciation is one of the fundamentals of language learning, and it is considered a primary factor of spoken language when it comes to an understanding and being understood by others. The persistent presence of high er…

Language Modellingspeech-recognitionSpeech Recognition

speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment

2021-04-03 · Junbo Zhang, Zhiwen Zhang, Yongqing Wang, Zhiyong Yan 외

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are c…

Phone-level pronunciation scoringSentencespeech-recognition

CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis

2021-11-16 · Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung 외

Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems. End-to-end (e2e) approaches are becoming dominant in MDD. However an e2e MDD model usual…

Multi-Task LearningPhone-level pronunciation scoring

Class LM and word mapping for contextual biasing in End-to-End ASR

2020-07-10 · Rongqing Huang, Ossama Abdel-hamid, Xinwei Li, Gunnar Evermann

In recent years, all-neural, end-to-end (E2E) ASR systems gained rapid interest in the speech recognition community. They convert speech input to text units in a single trainable Neural Network model. In ASR, many uttera…

speech-recognitionSpeech Recognition