paper-with-me

홈 › Papers

Improving Yorùbá Diacritic Restoration

2020-03-23 · Iroro Orife, David I. Adelani, Timi Fasubaa, Victor Williamson, Wuraola Fisayo Oyewusi, Olamilekan Wahab, Kola Tubosun

Yor\ub\'a is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any computational Speech or Natural Language Processing tasks. However diacritic marks are commonly excluded from electronic texts due to limited device and application support as well as general education on proper usage. We report on recent efforts at dataset cultivation. By aggregating and improving disparate texts from the web and various personal libraries, we were able to significantly grow our clean Yor\ub\'a dataset from a majority Bibilical text corpora with three sources to millions of tokens from over a dozen sources. We evaluate updated diacritic restoration models on a new, general purpose, public-domain Yor\ub\'a evaluation dataset of modern journalistic news text, selected to be multi-purpose and reflecting contemporary usage. All pre-trained models, datasets and source-code have been released as an open-source project to advance efforts on Yor\ub\'a language technology.

📄 PDF Abstract BibTeX arXiv:2003.10564

Code (1)

Niger-Volta-LTI/yoruba-adr 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Automatic Restoration of Diacritics for Speech Data Sets

2023-11-15 · Sara Shatnawi, Sawsan Alqahtani, Hanan Aldarmaki

Automatic text-based diacritic restoration models generally have high diacritic error rates when applied to speech transcripts as a result of domain and style shifts in spoken language. In this work, we explore the possi…

Lexical Disambiguation of Igbo using Diacritic Restoration

2017-04-01 · WS 2017 4 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe

Properly written texts in Igbo, a low-resource African language, are rich in both orthographic and tonal diacritics. Diacritics are essential in capturing the distinctions in pronunciation and meaning of words, as well a…

BIG-bench Machine LearningGeneral Classification

Igbo Diacritic Restoration using Embedding Models

2018-06-01 · NAACL 2018 6 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe, Enemouh Chioma

Igbo is a low-resource language spoken by approximately 30 million people worldwide. It is the native language of the Igbo people of south-eastern Nigeria. In Igbo language, diacritics - orthographic and tonal - play a h…

Machine TranslationWord Embeddings

Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study

2025-11-17 · Mihai Nadas, Laura Diosan arxiv

Automatic diacritic restoration is crucial for text processing in languages with rich diacritical marks, such as Romanian. This study evaluates the performance of several large language models (LLMs) in restoring diacrit…

A Multitask Learning Approach for Diacritic Restoration

2020-06-07 · ACL 2020 6 · Sawsan Alqahtani, Ajay Mishra, Mona Diab

In many languages like Arabic, diacritics are used to specify pronunciations as well as meanings. Such diacritics are often omitted in written text, increasing the number of possible pronunciations and meanings for a wor…

Multi-Task LearningPart-Of-Speech Tagging