Universal Dependencies for Albanian
In this paper, we introduce the first Universal Dependencies (UD) treebank for standard Albanian, consisting of 60 sentences collected from the Albanian Wikipedia, annotated with lemmas, universal part-of-speech tags, morphological features and syntactic dependencies. In addition to presenting the treebank itself, we discuss a selection of linguistic constructions in Albanian whose analysis in UD is not self-evident, including core arguments and the status of indirect objects, pronominal clitics, genitive constructions, prearticulated adjectives, and modal verbs.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models
In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…
LemmatizationMorphological TaggingPart-Of-Speech TaggingAlbanian Language Identification in Text Documents
In this work we investigate the accuracy of standard and state-of-the-art language identification methods in identifying Albanian in written text documents. A dataset consisting of news articles written in Albanian has b…
ArticlesGeneral ClassificationLanguage IdentificationAlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian
Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety e…
AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled …
Ensemble LearningAlbNER: A Corpus for Named Entity Recognition in Albanian
Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, …
Articlesnamed-entity-recognitionNamed Entity RecognitionNER