GIPFA: Generating IPA Pronunciation from Audio
Transcribing spoken audio samples into the International Phonetic Alphabet (IPA) has long been reserved for experts. In this study, we examine the use of an Artificial Neural Network (ANN) model to automatically extract the IPA phonemic pronunciation of a word based on its audio pronunciation, hence its name Generating IPA Pronunciation From Audio (GIPFA). Based on the French Wikimedia dictionary, we trained our model which then correctly predicted 75% of the IPA pronunciations tested. Interestingly, by studying inference errors, the model made it possible to highlight possible errors in the dataset as well as to identify the closest phonemes in French.
Code (1)
Similar Papers 제목 키워드 기반
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciat…
Knowledge DistillationCorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction
This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations…
speech-recognitionSpeech RecognitionAUTOMATIC PRONUNCIATION MISTAKE DETECTOR PROJECT REPORT
Given the drawbacks of traditional English pronunciation correction systems, such as failure to provide timely feedback and correct learners' pronunciation errors, slow improvement of learners' English proficiency, and e…
Mistake Detectionspeech-recognitionSpeech RecognitionDTW-SiameseNet: Dynamic Time Warped Siamese Network for Mispronunciation Detection and Correction
Personal Digital Assistants (PDAs) - such as Siri, Alexa and Google Assistant, to name a few - play an increasingly important role to access information and complete tasks spanning multiple domains, and by diverse groups…
Dynamic Time WarpingMetric LearningPrivacy PreservingSpeech Synthesis+3ASMDD: Arabic Speech Mispronunciation Detection Dataset
The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the A…