paper-with-me

Papers

EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

2024-03-28 · Atnafu Lambebo Tonja, Olga Kolesnikova, Alexander Gelbukh, Jugal Kalita

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for low-resource languages. This is due to the smaller size of available parallel corpora in these languages, if such corpora are available at all. NLP in Ethiopian languages suffers from the same issues due to the unavailability of publicly accessible datasets for NLP tasks, including MT. To help the research community and foster research for Ethiopian languages, we introduce EthioMT -- a new parallel corpus for 15 languages. We also create a new benchmark by collecting a dataset for better-researched languages in Ethiopia. We evaluate the newly collected corpus and the benchmark dataset for 23 Ethiopian languages using transformer and fine-tuning approaches.

📄 PDF Abstract BibTeX arXiv:2403.19365

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNews ClassificationQuestion Answering

Similar Papers 제목 키워드 기반

Parallel Corpora for bi-lingual English-Ethiopian Languages Statistical Machine Translation

2018-08-01 · COLING 2018 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge{'}ez. The corpora are used for conducting a bi…

Machine TranslationTranslation

Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo, and Wolaytta

2020-07-01 · WS 2020 7 · Solomon Teferra Abate, Martha Yifiru Tachbelie, Michael Melese, Hafte Abera 외

Automatic Speech Recognition (ASR) is one of the most important technologies to help people live a better life in the 21st century. However, its development requires a big speech corpus for a language. The development of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

English-Ethiopian Languages Statistical Machine Translation

2019-08-01 · WS 2019 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge{'}ez. The corpora are used for conducting bi-d…

Machine TranslationTranslation

Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo and Wolaytta

2020-05-01 · LREC 2020 5 · Solomon Teferra Abate, Martha Yifiru Tachbelie, Michael Melese, Hafte Abera 외

Automatic Speech Recognition (ASR) is one of the most important technologies to support spoken communication in modern life. However, its development benefits from large speech corpus. The development of such a corpus is…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Parallel Corpora for bi-Directional Statistical Machine Translation for Seven Ethiopian Language Pairs

2018-08-01 · COLING 2018 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe the development of parallel corpora for Ethiopian Languages: Amharic, Tigrigna, Afan-Oromo, Wolaytta and Geez. To check the usability of all the corpora we conducted baseline bi-directional sta…

Machine TranslationTranslation