Improved Text Language Identification for the South African Languages
Virtual assistants and text chatbots have recently been gaining popularity. Given the short message nature of text-based chat interactions, the language identification systems of these bots might only have 15 or 20 characters to make a prediction. However, accurate text language identification is important, especially in the early stages of many multilingual natural language processing pipelines. This paper investigates the use of a naive Bayes classifier, to accurately predict the language family that a piece of text belongs to, combined with a lexicon based classifier to distinguish the specific South African language that the text is written in. This approach leads to a 31% reduction in the language detection error. In the spirit of reproducible research the training and testing datasets as well as the code are published on github. Hopefully it will be useful to create a text language identification shared task for South African languages.
Code (1)
Tasks
Language IdentificationSimilar Papers 제목 키워드 기반
Short Text Language Identification for Under Resourced Languages
The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the …
Language IdentificationPreparing the Vuk'uzenzele and ZA-gov-multilingual South African multilingual corpora
This paper introduces two multilingual government themed corpora in various South African languages. The corpora were collected by gathering the South African Government newspaper (Vuk'uzenzele), as well as South African…
Language ModelingLanguage ModellingMachine TranslationNMT+1African WordNet: A Viable Tool for Sense Discrimination in the Indigenous African Languages of South Africa
In promoting a multilingual South Africa, the government is encouraging people to speak more than one language. In order to comply with this initiative, people choose to learn the languages which they do not speak as hom…
Benchmarking Neural Machine Translation for Southern African Languages
Unlike major Western languages, most African languages are very low-resourced. Furthermore, the resources that do exist are often scattered and difficult to obtain and discover. As a result, the data and code for existin…
BenchmarkingMachine TranslationTranslationA Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
This study investigates how cultural proximity affects the ability to detect AI-generated fake news by comparing South African participants with those from other nationalities. As large language models increasingly enabl…