Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text
Yor\ub\'a is a widely spoken West African language with a writing system
rich in tonal and orthographic diacritics. With very few exceptions, diacritics
are omitted from electronic texts, due to limited device and application
support. Diacritics provide morphological information, are crucial for lexical
disambiguation, pronunciation and are vital for any Yor\ub\'a text-to-speech
(TTS), automatic speech recognition (ASR) and natural language processing (NLP)
tasks. Reframing Automatic Diacritic Restoration (ADR) as a machine translation
task, we experiment with two different attentive Sequence-to-Sequence neural
models to process undiacritized text. On our evaluation dataset, this approach
produces diacritization error rates of less than 5%. We have released
pre-trained models, datasets and source-code as an open-source project to
advance efforts on Yor\`ub\'a language technology.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognitionSpeech Recognitiontext-to-speechText to SpeechTranslationSimilar Papers 제목 키워드 기반
Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration
Kashmiri, an Indo-Aryan language written in a modified Perso-Arabic script, frequently omits diacritic marks in digital text, creating ambiguity and challenging downstream NLP applications. We present Koshur Diacritizer,…
A System for Diacritizing Four Varieties of Arabic
Short vowels, aka diacritics, are more often omitted when writing different varieties of Arabic including Modern Standard Arabic (MSA), Classical Arabic (CA), and Dialectal Arabic (DA). However, diacritics are required t…
Feature Engineeringtext-to-speechText to SpeechEfficient Convolutional Neural Networks for Diacritic Restoration
Diacritic restoration has gained importance with the growing need for machines to understand written texts. The task is typically modeled as a sequence labeling problem and currently Bidirectional Long Short Term Memory …
Automatic Restoration of Diacritics for Speech Data Sets
Automatic text-based diacritic restoration models generally have high diacritic error rates when applied to speech transcripts as a result of domain and style shifts in spoken language. In this work, we explore the possi…
Corpus-Based Approaches to Igbo Diacritic Restoration
With natural language processing (NLP), researchers aim to enable computers to identify and understand patterns in human languages. This is often difficult because a language embeds many dynamic and varied properties in …