paper-with-me

Papers

Natural Language Processing in Ethiopian Languages: Current State, Challenges, and Opportunities

2023-03-25 · Atnafu Lambebo Tonja, Tadesse Destaw Belay, Israel Abebe Azime, Abinew Ali Ayele, Moges Ahmed Mehamed, Olga Kolesnikova, Seid Muhie Yimam

This survey delves into the current state of natural language processing (NLP) for four Ethiopian languages: Amharic, Afaan Oromo, Tigrinya, and Wolaytta. Through this paper, we identify key challenges and opportunities for NLP research in Ethiopia. Furthermore, we provide a centralized repository on GitHub that contains publicly available resources for various NLP tasks in these languages. This repository can be updated periodically with contributions from other researchers. Our objective is to identify research gaps and disseminate the information to NLP researchers interested in Ethiopian languages and encourage future research in this domain.

📄 PDF Abstract BibTeX arXiv:2303.14406

Code (1)

EthioNLP/Ethiopian-Language-Survey 공식 구현

Similar Papers 제목 키워드 기반

EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

2024-03-28 · Atnafu Lambebo Tonja, Olga Kolesnikova, Alexander Gelbukh, Jugal Kalita

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the per…

Machine TranslationNews ClassificationQuestion Answering

EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation

2024-03-20 · Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu 외

Large language models (LLMs) have gained popularity recently due to their outstanding performance in various downstream Natural Language Processing (NLP) tasks. However, low-resource languages are still lagging behind cu…

Diversity

Analysis of GlobalPhone and Ethiopian Languages Speech Corpora for Multilingual ASR

2020-05-01 · LREC 2020 5 · Martha Yifiru Tachbelie, Solomon Teferra Abate, Tanja Schultz

In this paper, we present the analysis of GlobalPhone (GP) and speech corpora of Ethiopian languages (Amharic, Tigrigna, Oromo and Wolaytta). The aim of the analysis is to select speech data from GP for the development o…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Parallel Corpora for bi-lingual English-Ethiopian Languages Statistical Machine Translation

2018-08-01 · COLING 2018 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge{'}ez. The corpora are used for conducting a bi…

Machine TranslationTranslation

English-Ethiopian Languages Statistical Machine Translation

2019-08-01 · WS 2019 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge{'}ez. The corpora are used for conducting bi-d…

Machine TranslationTranslation