Investigating an approach for low resource language dataset creation, curation and classification: Setswana and Sepedi
The recent advances in Natural Language Processing have been a boon for well-represented languages in terms of available curated data and research resources. One of the challenges for low-resourced languages is clear guidelines on the collection, curation and preparation of datasets for different use-cases. In this work, we take on the task of creation of two datasets that are focused on news headlines (i.e short text) for Setswana and Sepedi and creation of a news topic classification task. We document our work and also present baselines for classification. We investigate an approach on data augmentation, better suited to low resource languages, to improve the performance of the classifiers
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationData AugmentationGeneral ClassificationTopic ClassificationSimilar Papers 제목 키워드 기반
Low resource language dataset creation, curation and classification: Setswana and Sepedi -- Extended Abstract
The recent advances in Natural Language Processing have only been a boon for well represented languages, negating research in lesser known global languages. This is in part due to the availability of curated data and res…
Data AugmentationGeneral ClassificationTopic ClassificationAI4D -- African Language Program
Advances in speech and language technologies enable tools such as voice-search, text-to-speech, speech recognition and machine translation. These are however only available for high resource languages like English, Frenc…
Machine Translationspeech-recognitionSpeech Recognitiontext-to-speech+2AI4D - African Language Dataset Challenge
As language and speech technologies become more advanced, the lack of fundamental digital resources for African languages, such as data, spell checkers and PoS taggers, means that the digital divide between these languag…
POSNo Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem
Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English,…
The evolving infrastructure for language resources and the role for data scientists
In the context of ongoing developments as regards the creation of a sustainable, interoperable language resource infrastructure and spreading ideas of the need for open access, not only of research publications but also …