paper-with-me

Papers

NorNE: Annotating Named Entities for Norwegian

2019-11-27 · LREC 2020 5 · Fredrik Jørgensen, Tobias Aasmoe, Anne-Stine Ruud Husevåg, Lilja Øvrelid, Erik Velldal

This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokm{\aa}l and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture.

📄 PDF Abstract BibTeX arXiv:1911.12146

Code (1)

ltgoslo/norne 공식 구현

Similar Papers 제목 키워드 기반

Aligning the Norwegian UD Treebank with Entity and Coreference Information

2023-05-22 · Tollef Emil Jørgensen, Andre Kåsen

This paper presents a merged collection of entity and coreference annotated data grounded in the Universal Dependencies (UD) treebanks for the two written forms of Norwegian: Bokm{\aa}l and Nynorsk. The aligned and conve…

DaNE: A Named Entity Resource for Danish

2020-05-01 · LREC 2020 5 · Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted 외

We present a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme: DaNE. It is the largest publicly available, Danish named entity gold annotation. We evaluate the…

Cross-Lingual Transfernamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Annotating named entities in clinical text by combining pre-annotation and active learning

2013-08-01 · ACL 2013 8 · Maria Skeppstedt
Active Learning

Annotating Named Entities in Consumer Health Questions

2016-05-01 · LREC 2016 5 · Halil Kilicoglu, Asma Ben Abacha, Yassine Mrabet, Kirk Roberts 외

We describe a corpus of consumer health questions annotated with named entities. The corpus consists of 1548 de-identified questions about diseases and drugs, written in English. We defined 15 broad categories of biomedi…

Annotating evaluative sentences for sentiment analysis: a dataset for Norwegian

2019-09-01 · WS (NoDaLiDa) 2019 9 · Petter Mæhlum, Jeremy Barnes, Lilja Øvrelid, Erik Velldal

This paper documents the creation of a large-scale dataset of evaluative sentences – i.e. both subjective and objective sentences that are found to be sentiment-bearing – based on mixed-domain professional reviews from v…

Sentiment Analysis