paper-with-me

홈 › Papers

Generating Varied Training Corpora in Runyankore Using a Combined Semantic and Syntactic, Pattern-Grammar-based Approach

2020-12-01 · INLG (ACL) 2020 12 · Joan Byamugisha

Machine learning algorithms have been applied to achieve high levels of accuracy in tasks associated with the processing of natural language. However, these algorithms require large amounts of training data in order to perform efficiently. Since most Bantu languages lack the required training corpora because they are computationally under-resourced, we investigated how to generate a large varied training corpus in Runyankore, a Bantu language indigenous to Uganda. We found the use of a combined semantic and syntactic, pattern and grammar-based approach to be applicable to this purpose, and used it to generate one million sentences, both labelled and unlabelled, which can be applied as training data for machine learning algorithms. The generated text was evaluated in two ways: (1) assessing the semantics encoded in word embeddings obtained from the generated text, which showed correct word similarity; and (2) applying the labelled data to tasks such as sentiment analysis, which achieved satisfactory levels of accuracy.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningSentiment AnalysisWord EmbeddingsWord Similarity

Similar Papers 제목 키워드 기반

Noun Class Disambiguation in Runyankore and Related Languages

2022-10-01 · COLING 2022 10 · Joan Byamugisha

Bantu languages are spoken by communities in more than half of the countries on the African continent by an estimated third of a billion people. Despite this populous and the amount of high quality linguistic research do…

Towards Computational Resource Grammars for Runyankore and Rukiga

2020-05-01 · LREC 2020 5 · David Bamutura, Peter Ljungl{\"o}f, Peter Nebende

In this paper, we present computational resource grammars of Runyankore and Rukiga (R{\&}R) languages. Runyankore and Rukiga are two under-resourced Bantu Languages spoken by about 6 million people indigenous to South- W…

Descriptive

Tense and Aspect in Runyankore Using a Context-Free Grammar

2016-09-01 · WS 2016 9 · Joan Byamugisha, C. Maria Keet, Brian DeRenzi
Text Generation

Towards a Resource Grammar for Runyankore and Rukiga

2019-08-01 · WS 2019 8 · David Bamutura, Peter Ljungl{\"o}f

Currently, there is a lack of computational grammar resources for many under-resourced languages which limits the ability to develop Natural Language Processing (NLP) tools and applications such as Multilingual Document …

Machine TranslationTranslation

Toward an NLG System for Bantu languages: first steps with Runyankore (demo)

2017-09-01 · WS 2017 9 · Joan Byamugisha, C. Maria Keet, Brian DeRenzi

There are many domain-specific and language-specific NLG systems, of which it may be possible to adapt to related domains and languages. The languages in the Bantu language family have their own set of features distinct …

Text Generation