paper-with-me

Papers

Building Multilingual Corpora for a Complex Named Entity Recognition and Classification Hierarchy using Wikipedia and DBpedia

2022-12-14 · Diego Alves, Gaurish Thakkar, Gabriel Amaral, Tin Kuculo, Marko Tadić

With the ever-growing popularity of the field of NLP, the demand for datasets in low resourced-languages follows suit. Following a previously established framework, in this paper, we present the UNER dataset, a multilingual and hierarchical parallel corpus annotated for named-entities. We describe in detail the developed procedure necessary to create this type of dataset in any language available on Wikipedia with DBpedia information. The three-step procedure extracts entities from Wikipedia articles, links them to DBpedia, and maps the DBpedia sets of classes to the UNER labels. This is followed by a post-processing procedure that significantly increases the number of identified entities in the final results. The paper concludes with a statistical and qualitative analysis of the resulting dataset.

📄 PDF Abstract BibTeX arXiv:2212.07429

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesNamed Entity RecognitionNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

UM6P-CS at SemEval-2022 Task 11: Enhancing Multilingual and Code-Mixed Complex Named Entity Recognition via Pseudo Labels using Multilingual Transformer

2022-04-28 · SemEval (NAACL) 2022 7 · Abdellah El Mekki, Abdelkader El Mahdaouy, Mohammed Akallouch, Ismail Berrada 외

Building real-world complex Named Entity Recognition (NER) systems is a challenging task. This is due to the complexity and ambiguity of named entities that appear in various contexts such as short input sentences, emerg…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2

Language-Independent Named Entity Analysis Using Parallel Projection and Rule-Based Disambiguation

2017-04-01 · WS 2017 4 · James Mayfield, Paul McNamee, Cash Costello

The 2017 shared task at the Balto-Slavic NLP workshop requires identifying coarse-grained named entities in seven languages, identifying each entity{'}s base form, and clustering name mentions across the multilingual set…

Clusteringnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

SemEval-2022 Task 11: Multilingual Complex Named Entity Recognition (MultiCoNER)

2022-07-01 · SemEval (NAACL) 2022 7 · Shervin Malmasi, Anjie Fang, Besnik Fetahu, Sudipta Kar 외

We present the findings of SemEval-2022 Task 11 on Multilingual Complex Named Entity Recognition MULTICONER. Divided into 13 tracks, the task focused on methods to identify complex named entities (like names of movies, p…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark

2026-04-14 · Terra Blevins, Stephen Mayhew, Marek Šuppa, Hila Gonen 외 arxiv

While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these assumptions remain scarce. The Universal …

MultiCoNER: A Large-scale Multilingual dataset for Complex Named Entity Recognition

2022-08-30 · COLING 2022 10 · Shervin Malmasi, Anjie Fang, Besnik Fetahu, Sudipta Kar 외

We present MultiCoNER, a large multilingual dataset for Named Entity Recognition that covers 3 domains (Wiki sentences, questions, and search queries) across 11 languages, as well as multilingual and code-mixing subsets.…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2