paper-with-me

Papers

Universal Dependencies for Albanian

2020-12-01 · UDW (COLING) 2020 12 · Marsida Toska, Joakim Nivre, Daniel Zeman

In this paper, we introduce the first Universal Dependencies (UD) treebank for standard Albanian, consisting of 60 sentences collected from the Albanian Wikipedia, annotated with lemmas, universal part-of-speech tags, morphological features and syntactic dependencies. In addition to presenting the treebank itself, we discuss a selection of linguistic constructions in Albanian whose analysis in UD is not self-evident, including core arguments and the status of indirect objects, pronominal clitics, genitive constructions, prearticulated adjectives, and modal verbs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

2019-12-02 · Nelda Kote, Marenglen Biba, Jenna Kanerva, Samuel Rönnqvist 외

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…

LemmatizationMorphological TaggingPart-Of-Speech Tagging

Albanian Language Identification in Text Documents

2019-01-14 · Klesti Hoxha, Artur Baxhaku

In this work we investigate the accuracy of standard and state-of-the-art language identification methods in identifying Albanian in written text documents. A dataset consisting of news articles written in Albanian has b…

ArticlesGeneral ClassificationLanguage Identification

AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

2026-05-26 · Wajdi Zaghouani, Kholoud K. Aldous, Isra Fejzullaj arxiv

Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved. We present AlbanianLLMSafety, the first publicly available safety e…

AlbNews: A Corpus of Headlines for Topic Modeling in Albanian

2024-02-06 · Erion Çano, Dario Lamaj

The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled …

Ensemble Learning

AlbNER: A Corpus for Named Entity Recognition in Albanian

2023-09-15 · Erion Çano

Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, …

Articlesnamed-entity-recognitionNamed Entity RecognitionNER