AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled news headlines and 2600 unlabeled ones in Albanian. The data can be freely used for conducting topic modeling research. We report the initial classification scores of some traditional machine learning classifiers trained with the AlbNews samples. These results show that basic models outrun the ensemble learning ones and can serve as a baseline for future experiments.
Code (1)
Tasks
Ensemble LearningSimilar Papers 제목 키워드 기반
AlbNER: A Corpus for Named Entity Recognition in Albanian
Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, …
Articlesnamed-entity-recognitionNamed Entity RecognitionNERMorphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models
In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…
LemmatizationMorphological TaggingPart-Of-Speech TaggingAlbanian Language Identification in Text Documents
In this work we investigate the accuracy of standard and state-of-the-art language identification methods in identifying Albanian in written text documents. A dataset consisting of news articles written in Albanian has b…
ArticlesGeneral ClassificationLanguage IdentificationLearning to Flip the Bias of News Headlines
This paper introduces the task of {``}flipping{''} the bias of news articles: Given an article with a political bias (left or right), generate an article with the same topic but opposite bias. To study this task, we crea…
ArticlesText GenerationAlbMoRe: A Corpus of Movie Reviews for Sentiment Analysis in Albanian
Lack of available resources such as text corpora for low-resource languages seriously hinders research on natural language processing and computational linguistics. This paper presents AlbMoRe, a corpus of 800 sentiment …
Sentiment Analysis