paper-with-me

Papers

DANSK and DaCy 2.6.0: Domain Generalization of Danish Named Entity Recognition

2024-02-28 · Kenneth Enevoldsen, Emil Trenckner Jessen, Rebekah Baglini

Named entity recognition is one of the cornerstones of Danish NLP, essential for language technology applications within both industry and research. However, Danish NER is inhibited by a lack of available datasets. As a consequence, no current models are capable of fine-grained named entity recognition, nor have they been evaluated for potential generalizability issues across datasets and domains. To alleviate these limitations, this paper introduces: 1) DANSK: a named entity dataset providing for high-granularity tagging as well as within-domain evaluation of models across a diverse set of domains; 2) DaCy 2.6.0 that includes three generalizable models with fine-grained annotation; and 3) an evaluation of current state-of-the-art models' ability to generalize across domains. The evaluation of existing and new models revealed notable performance discrepancies across domains, which should be addressed within the field. Shortcomings of the annotation quality of the dataset and its impact on model training and evaluation are also discussed. Despite these limitations, we advocate for the use of the new dataset DANSK alongside further work on the generalizability within Danish NER.

📄 PDF Abstract BibTeX arXiv:2402.18209

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalizationnamed-entity-recognitionNamed Entity RecognitionNER

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

DaCy: A Unified Framework for Danish NLP

2021-07-12 · Kenneth Enevoldsen, Lasse Hansen, Kristoffer Nielbo

Danish natural language processing (NLP) has in recent years obtained considerable improvements with the addition of multiple new datasets and models. However, at present, there is no coherent framework for applying stat…

Data AugmentationDependency Parsingnamed-entity-recognitionNamed Entity Recognition+2

The Lacunae of Danish Natural Language Processing

2019-09-01 · WS (NoDaLiDa) 2019 9 · Andreas Kirkedal, Barbara Plank, Leon Derczynski, Natalie Schluter

Danish is a North Germanic language spoken principally in Denmark, a country with a long tradition of technological and scientific innovation. However, the language has received relatively little attention from a technol…

Compiling a Suitable Level of Sense Granularity in a Lexicon for AI Purposes: The Open Source COR Lexicon

2022-06-01 · LREC 2022 6 · Bolette Pedersen, Nathalie Carmen Hau Sørensen, Sanni Nimb, Ida Flørke 외

We present The Central Word Register for Danish (COR), which is an open source lexicon project for general AI purposes funded and initiated by the Danish Agency for Digitisation as part of an AI initiative embarked by th…

DaNE: A Named Entity Resource for Danish

2020-05-01 · LREC 2020 5 · Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted 외

We present a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme: DaNE. It is the largest publicly available, Danish named entity gold annotation. We evaluate the…

Cross-Lingual Transfernamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

DaN+: Danish Nested Named Entities and Lexical Normalization

2021-05-24 · COLING 2020 8 · Barbara Plank, Kristian Nørgaard Jensen, Rob van der Goot

This paper introduces DaN+, a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resou…

Cross-Lingual TransferLexical NormalizationMulti-Task Learningnamed-entity-recognition+3