The Challenge of Diacritics in Yoruba Embeddings
The major contributions of this work include the empirical establishment of a better performance for Yoruba embeddings from undiacritized (normalized) dataset and provision of new analogy sets for evaluation. The Yoruba language, being a tonal language, utilizes diacritics (tonal marks) in written form. We show that this affects embedding performance by creating embeddings from exactly the same Wikipedia dataset but with the second one normalized to be undiacritized. We further compare average intrinsic performance with two other work (using analogy test set & WordSim) and we obtain the best performance in WordSim and corresponding Spearman correlation.
Code (1)
Similar Papers 제목 키워드 기반
The Effect of Domain and Diacritics in Yoruba–English Neural Machine Translation
Massively multilingual machine translation (MT) has shown impressive capabilities and including zero and few-shot translation between low-resource language pairs. However and these models are often evaluated on high-reso…
BenchmarkingMachine TranslationTranslationEmbed More Ignore Less (EMIL): Exploiting Enriched Representations for Arabic NLP
Our research focuses on the potential improvements of exploiting language specific characteristics in the form of embeddings by neural networks. More specifically, we investigate the capability of neural techniques and e…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3Diacritics Restoration using BERT with Analysis on Czech language
We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduct a detailed error analysis on Czech, a …
Croatian Text DiacritizationCzech Text DiacritizationFrench Text DiacritizationHungarian Text Diacritization+8A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour
We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text …
Igbo Diacritic Restoration using Embedding Models
Igbo is a low-resource language spoken by approximately 30 million people worldwide. It is the native language of the Igbo people of south-eastern Nigeria. In Igbo language, diacritics - orthographic and tonal - play a h…
Machine TranslationWord Embeddings