paper-with-me

홈 › Papers

KINNEWS and KIRNEWS: Benchmarking Cross-Lingual Text Classification for Kinyarwanda and Kirundi

2020-10-23 · COLING 2020 8 · Rubungo Andre Niyongabo, Hong Qu, Julia Kreutzer, Li Huang

Recent progress in text classification has been focused on high-resource languages such as English and Chinese. For low-resource languages, amongst them most African languages, the lack of well-annotated data and effective preprocessing, is hindering the progress and the transfer of successful methods. In this paper, we introduce two news datasets (KINNEWS and KIRNEWS) for multi-class classification of news articles in Kinyarwanda and Kirundi, two low-resource African languages. The two languages are mutually intelligible, but while Kinyarwanda has been studied in Natural Language Processing (NLP) to some extent, this work constitutes the first study on Kirundi. Along with the datasets, we provide statistics, guidelines for preprocessing, and monolingual and cross-lingual baseline models. Our experiments show that training embeddings on the relatively higher-resourced Kinyarwanda yields successful cross-lingual transfer to Kirundi. In addition, the design of the created datasets allows for a wider use in NLP beyond text classification in future studies, such as representation learning, cross-lingual learning with more distant languages, or as base for new annotations for tasks such as parsing, POS tagging, and NER. The datasets, stopwords, and pre-trained embeddings are publicly available at https://github.com/Andrews2017/KINNEWS-and-KIRNEWS-Corpus .

📄 PDF Abstract BibTeX arXiv:2010.12174

Code (1)

Andrews2017/KINNEWS-and-KIRNEWS-Corpus 공식 구현 pytorch

Tasks

ArticlesBenchmarkingCross-Lingual TransferGeneral ClassificationMulti-class ClassificationNERNews ClassificationPOSPOS TaggingRepresentation Learningtext-classificationText Classification

Similar Papers 제목 키워드 기반

MTG: A Benchmarking Suite for Multilingual Text Generation

2021-10-16 · ACL ARR October 2021 10 · Anonymous

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first and largest multilingual multiway text generation benchmark with 400k human-annotated data for four tasks …

BenchmarkingQuestion GenerationQuestion-GenerationStory Generation+3

Benchmarking Cross-Lingual Semantic Alignment in Multilingual Embeddings

2025-12-29 · Wen G. Gong arxiv

With hundreds of multilingual embedding models available, practitioners lack clear guidance on which provide genuine cross-lingual semantic alignment versus task performance through language-specific patterns. Task-drive…

GPTs and Language Barrier: A Cross-Lingual Legal QA Examination

2024-03-26 · Ha-Thanh Nguyen, Hiroaki Yamada, Ken Satoh

In this paper, we explore the application of Generative Pre-trained Transformers (GPTs) in cross-lingual legal Question-Answering (QA) systems using the COLIEE Task 4 dataset. In the COLIEE Task 4, given a statement and …

ArticlesBenchmarkingQuestion Answeringvalid

Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval

2023-11-03 · Jinrui Yang, Timothy Baldwin, Trevor Cohn

We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a mult…

BenchmarkingFairnessInformation RetrievalRetrieval+1

BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer

2023-05-24 · Akari Asai, Sneha Kudugunta, Xinyan Velocity Yu, Terra Blevins 외

Despite remarkable advancements in few-shot generalization in natural language processing, most models are developed and evaluated primarily in English. To facilitate research on few-shot cross-lingual transfer, we intro…

BenchmarkingCross-Lingual TransferIn-Context Learning