paper-with-me

Papers

Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic

2024-08-05 · Yassine El Kheir, Hamdy Mubarak, Ahmed Ali, Shammur Absar Chowdhury

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic sound sets. The proposed framework utilized a quantized sequence of input with(out) continuous pretrained self-supervised representation. We show the efficacy of the pipeline using limited data for Arabic, a dialect-rich language containing more than 22 major dialects. Phonetically correct transcribed speech resources for dialectal Arabic are scarce. Therefore, we introduce ArabVoice15, a first-of-its-kind, curated test set featuring 5 hours of dialectal speech across 15 Arab countries, with phonetically accurate transcriptions, including borrowed and dialect-specific sounds. We described in detail the annotation guideline along with the analysis of the dialectal confusion pairs. Our extensive evaluation includes both subjective -- human perception tests and objective measures. Our empirical results, reported with three test sets, show that with only one and half hours of training data, our model improve character error rate by ~ 7\% in ArabVoice15 compared to the baseline.

📄 PDF Abstract BibTeX arXiv:2408.02430

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Using Ambiguity Detection to Streamline Linguistic Annotation

2016-12-01 · WS 2016 12 · Wajdi Zaghouani, Abdelati Hawwari, Sawsan Alqahtani, Houda Bouamor 외

Arabic writing is typically underspecified for short vowels and other markups, referred to as diacritics. In addition to the lexical ambiguity exhibited in most languages, the lack of diacritics in written Arabic adds an…

Automatic Speech Recognition (ASR)Machine TranslationSpeech Recognition

Recovering Missing Characters in Old Hawaiian Writing

2018-10-01 · EMNLP 2018 10 · Brendan Shillingford, Oiwi Parker Jones

In contrast to the older writing system of the 19th century, modern Hawaiian orthography employs characters for long vowels and glottal stops. These extra characters account for about one-third of the phonemes in Hawaiia…

Language ModelingLanguage ModellingReading ComprehensionTransliteration

Composing RNNs and FSTs for Small Data: Recovering Missing Characters in Old Hawaiian Text

2022-07-24 · Oiwi Parker Jones, Brendan Shillingford

In contrast to the older writing system of the 19th century, modern Hawaiian orthography employs characters for long vowels and glottal stops. These extra characters account for about one-third of the phonemes in Hawaiia…

Reading ComprehensionTransliteration

Composing RNNs and FSTs for Small Data: Recovering Missing Characters in Old Hawaiian Text

2018-10-15 · NIPS Workshop IRASL 2018 · Anonymous

In contrast to the older writing system of the 19th century, modern Hawaiian orthography employs characters for long vowels and glottal stops. These extra characters account for about one-third of the phonemes in Hawaiia…

Deep LearningReading ComprehensionTransliteration

Highly Effective Arabic Diacritization using Sequence to Sequence Modeling

2019-06-01 · NAACL 2019 6 · Hamdy Mubarak, Ahmed Abdelali, Hassan Sajjad, Younes Samih 외

Arabic text is typically written without short vowels (or diacritics). However, their presence is required for properly verbalizing Arabic and is hence essential for applications such as text to speech. There are two typ…

Feature EngineeringMachine Translationtext-to-speechText to Speech+1