paper-with-me

홈 › Papers

GIPFA: Generating IPA Pronunciation from Audio

2020-06-13 · Xavier Marjou

Transcribing spoken audio samples into the International Phonetic Alphabet (IPA) has long been reserved for experts. In this study, we examine the use of an Artificial Neural Network (ANN) model to automatically extract the IPA phonemic pronunciation of a word based on its audio pronunciation, hence its name Generating IPA Pronunciation From Audio (GIPFA). Based on the French Wikimedia dictionary, we trained our model which then correctly predicted 75% of the IPA pronunciations tested. Interestingly, by studying inference errors, the model made it possible to highlight possible errors in the dataset as well as to identify the closest phonemes in French.

📄 PDF Abstract BibTeX arXiv:2006.07573

Code (1)

marxav/gipfa 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS

2024-10-19 · Tuan Nam Nguyen, Seymanur Aktı, Ngoc Quan Pham, Alexander Waibel

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciat…

Knowledge Distillation

CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction

2022-04-12 · Daxin Tan, Liqun Deng, Nianzu Zheng, Yu Ting Yeung 외

This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations…

speech-recognitionSpeech Recognition

AUTOMATIC PRONUNCIATION MISTAKE DETECTOR PROJECT REPORT

2025-06-25 · ResearchGate 2025 6 · Kamal Acharya

Given the drawbacks of traditional English pronunciation correction systems, such as failure to provide timely feedback and correct learners' pronunciation errors, slow improvement of learners' English proficiency, and e…

Mistake Detectionspeech-recognitionSpeech Recognition

DTW-SiameseNet: Dynamic Time Warped Siamese Network for Mispronunciation Detection and Correction

2023-03-01 · Raviteja Anantha, Kriti Bhasin, Daniela de la Parra Aguilar, Prabal Vashisht 외

Personal Digital Assistants (PDAs) - such as Siri, Alexa and Google Assistant, to name a few - play an increasingly important role to access information and complete tasks spanning multiple domains, and by diverse groups…

Dynamic Time WarpingMetric LearningPrivacy PreservingSpeech Synthesis+3

ASMDD: Arabic Speech Mispronunciation Detection Dataset

2021-11-01 · Salah A. Aly, Abdelrahman Salah, Hesham M. Eraqi

The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the A…