paper-with-me

홈 › Papers

The Challenge of Diacritics in Yoruba Embeddings

2020-11-15 · Tosin P. Adewumi, Foteini Liwicki, Marcus Liwicki

The major contributions of this work include the empirical establishment of a better performance for Yoruba embeddings from undiacritized (normalized) dataset and provision of new analogy sets for evaluation. The Yoruba language, being a tonal language, utilizes diacritics (tonal marks) in written form. We show that this affects embedding performance by creating embeddings from exactly the same Wikipedia dataset but with the second one normalized to be undiacritized. We further compare average intrinsic performance with two other work (using analogy test set & WordSim) and we obtain the best performance in WordSim and corresponding Spearman correlation.

📄 PDF Abstract BibTeX arXiv:2011.07605

Code (1)

tosingithub/ydesk 공식 구현

Similar Papers 제목 키워드 기반

The Effect of Domain and Diacritics in Yoruba–English Neural Machine Translation

2021-08-01 · MTSummit 2021 8 · David Adelani, Dana Ruiter, Jesujoba Alabi, Damilola Adebonojo 외

Massively multilingual machine translation (MT) has shown impressive capabilities and including zero and few-shot translation between low-resource language pairs. However and these models are often evaluated on high-reso…

BenchmarkingMachine TranslationTranslation

Embed More Ignore Less (EMIL): Exploiting Enriched Representations for Arabic NLP

2020-12-01 · COLING (WANLP) 2020 12 · Ahmed Younes, Julie Weeds

Our research focuses on the potential improvements of exploiting language specific characteristics in the form of embeddings by neural networks. More specifically, we investigate the capability of neural techniques and e…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3

Diacritics Restoration using BERT with Analysis on Czech language

2021-05-24 · Jakub Náplava, Milan Straka, Jana Straková

We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduct a detailed error analysis on Czech, a …

Croatian Text DiacritizationCzech Text DiacritizationFrench Text DiacritizationHungarian Text Diacritization+8

A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

2026-07-17 · Kola Tubosun, Adedayo Oluokun, Hafiz Adewuyi, Dadepo Aderemi arxiv

We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text …

Igbo Diacritic Restoration using Embedding Models

2018-06-01 · NAACL 2018 6 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe, Enemouh Chioma

Igbo is a low-resource language spoken by approximately 30 million people worldwide. It is the native language of the Igbo people of south-eastern Nigeria. In Igbo language, diacritics - orthographic and tonal - play a h…

Machine TranslationWord Embeddings