paper-with-me

Papers

AlbNER: A Corpus for Named Entity Recognition in Albanian

2023-09-15 · Erion Çano

Scarcity of resources such as annotated text corpora for under-resourced languages like Albanian is a serious impediment in computational linguistics and natural language processing research. This paper presents AlbNER, a corpus of 900 sentences with labeled named entities, collected from Albanian Wikipedia articles. Preliminary results with BERT and RoBERTa variants fine-tuned and tested with AlbNER data indicate that model size has slight impact on NER performance, whereas language transfer has a significant one. AlbNER corpus and these obtained results should serve as baselines for future experiments.

📄 PDF Abstract BibTeX arXiv:2309.08741

Code (0)

등록된 구현이 없습니다.

Tasks

Articlesnamed-entity-recognitionNamed Entity RecognitionNER

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Residual Connection 설명 없음
Adam 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Weight Decay 설명 없음
Multi-Head Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

2019-12-02 · Nelda Kote, Marenglen Biba, Jenna Kanerva, Samuel Rönnqvist 외

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…

LemmatizationMorphological TaggingPart-Of-Speech Tagging

Albanian Language Identification in Text Documents

2019-01-14 · Klesti Hoxha, Artur Baxhaku

In this work we investigate the accuracy of standard and state-of-the-art language identification methods in identifying Albanian in written text documents. A dataset consisting of news articles written in Albanian has b…

ArticlesGeneral ClassificationLanguage Identification

pioNER: Datasets and Baselines for Armenian Named Entity Recognition

2018-10-19 · Tsolak Ghukasyan, Garnik Davtyan, Karen Avetisyan, Ivan Andrianov

In this work, we tackle the problem of Armenian named entity recognition, providing silver- and gold-standard datasets as well as establishing baseline results on popular models. We present a 163000-token named entity co…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings

UNER: Universal Named-Entity RecognitionFramework

2020-10-23 · Diego Alves, Tin Kuculo, Gabriel Amaral, Gaurish Thakkar 외

We introduce the Universal Named-Entity Recognition (UNER)framework, a 4-level classification hierarchy, and the methodology that isbeing adopted to create the first multilingual UNER corpus: the SETimesparallel corpus a…

Knowledge Graphsnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Named Entity Recognition in Information Security Domain for Russian

2019-09-01 · RANLP 2019 9 · Anastasiia Sirotina, Natalia Loukachevitch

In this paper we discuss the named entity recognition task for Russian texts related to cybersecurity. First of all, we describe the problems that arise in course of labeling unstructured texts from information security …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)