paper-with-me

Papers

A Sentiment Corpus for South African Under-Resourced Languages in a Multilingual Context

2022-06-01 · SIGUL (LREC) 2022 6 · Ronny Mabokela, Tim Schlippe

Multilingual sentiment analysis is a process of detecting and classifying sentiment based on textual information written in multiple languages. There has been tremendous research advancement on high-resourced languages such as English. However, progress on under-resourced languages remains underrepresented with limited opportunities for further development of natural language processing (NLP) technologies. Sentiment analysis (SA) for under-resourced language still is a skewed research area. Although, there are some considerable efforts in emerging African countries to develop such resources for under-resourced languages, languages such as indigenous South African languages still suffer from a lack of datasets. To the best of our knowledge, there is currently no dataset dedicated to SA research for South African languages in a multilingual context, i.e. comments are in different languages and may contain code-switching. In this paper, we present the first subset of the multilingual sentiment corpus SAfriSenti for the three most widely spoken languages in South Africa—English, Sepedi (i.e. Northern Sotho), and Setswana. This subset consists of over 40,000 annotated tweets in all the three languages including even 36.6% of code-switched texts. We present data collection, cleaning and annotation strategies that were followed to curate the dataset for these languages. Furthermore, we describe how we developed language-specific sentiment lexicons, morpheme-based sentiment taggers, conduct linguistic analyses and present possible solutions for the challenges of this sentiment dataset. We will release the dataset and sentiment lexicons to the research communities to advance the NLP research of under-resourced languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Similar Papers 제목 키워드 기반

Short Text Language Identification for Under Resourced Languages

2019-11-18 · Bernardt Duvenhage

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the …

Language Identification

TriLex: A Framework for Multilingual Sentiment Analysis in Low-Resource South African Languages

2025-12-02 · Mike Nkongolo, Hilton Vorster, Josh Warren, Trevor Naick 외 arxiv

Low-resource African languages remain underrepresented in sentiment analysis, limiting both lexical coverage and the performance of multilingual Natural Language Processing (NLP) systems. This study proposes TriLex, a th…

Sentiment Analysis

Benchmarking Neural Machine Translation for Southern African Languages

2019-06-17 · WS 2019 8 · Laura Martinus, Jade Z. Abbott

Unlike major Western languages, most African languages are very low-resourced. Furthermore, the resources that do exist are often scattered and difficult to obtain and discover. As a result, the data and code for existin…

BenchmarkingMachine TranslationTranslation

Semi-supervised learning approaches for predicting South African political sentiment for local government elections

2022-05-04 · Mashadi Ledwaba, Vukosi Marivate

This study aims to understand the South African political context by analysing the sentiments shared on Twitter during the local government elections. An emphasis on the analysis was placed on understanding the discussio…

Thinking globally, acting locally – Progress in the African Wordnet Project

2019-07-01 · GWC 2019 7 · Marissa Griesel, Sonja Bosch, Mampaka Lydia Mojapelo

The African Wordnet Project (AWN) includes all nine indigenous South African languages, namely isiZulu, isiXhosa, Setswana, Sesotho sa Leboa, Tshivenda, Siswati, Sesotho, isiNdebele and Xitsonga. The AWN currently includ…