paper-with-me

Papers

Investigating an approach for low resource language dataset creation, curation and classification: Setswana and Sepedi

2020-02-18 · LREC 2020 5 · Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya, Tumisho Mokgonyane, Rethabile Mokoena, Abiodun Modupe

The recent advances in Natural Language Processing have been a boon for well-represented languages in terms of available curated data and research resources. One of the challenges for low-resourced languages is clear guidelines on the collection, curation and preparation of datasets for different use-cases. In this work, we take on the task of creation of two datasets that are focused on news headlines (i.e short text) for Setswana and Sepedi and creation of a news topic classification task. We document our work and also present baselines for classification. We investigate an approach on data augmentation, better suited to low resource languages, to improve the performance of the classifiers

📄 PDF Abstract BibTeX arXiv:2003.04986

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationData AugmentationGeneral ClassificationTopic Classification

Similar Papers 제목 키워드 기반

Low resource language dataset creation, curation and classification: Setswana and Sepedi -- Extended Abstract

2020-03-30 · Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya 외

The recent advances in Natural Language Processing have only been a boon for well represented languages, negating research in lesser known global languages. This is in part due to the availability of curated data and res…

Data AugmentationGeneral ClassificationTopic Classification

AI4D -- African Language Program

2021-04-06 · Kathleen Siminyu, Godson Kalipe, Davor Orlic, Jade Abbott 외

Advances in speech and language technologies enable tools such as voice-search, text-to-speech, speech recognition and machine translation. These are however only available for high resource languages like English, Frenc…

Machine Translationspeech-recognitionSpeech Recognitiontext-to-speech+2

AI4D - African Language Dataset Challenge

2020-07-01 · WS 2020 7 · Kathleen Siminyu, Sackey Freshia

As language and speech technologies become more advanced, the lack of fundamental digital resources for African languages, such as data, spell checkers and PoS taggers, means that the digital divide between these languag…

POS

No Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem

2025-07-06 · Dasol Choi, Woomyoung Park, Youngsook Song arxiv

Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English,…

The evolving infrastructure for language resources and the role for data scientists

2014-05-01 · LREC 2014 5 · Nelleke Oostdijk, Henk van den Heuvel

In the context of ongoing developments as regards the creation of a sustainable, interoperable language resource infrastructure and spreading ideas of the need for open access, not only of research publications but also …