paper-with-me

홈 › Papers

Enhanced Entity Annotations for Multilingual Corpora

2022-06-01 · LREC 2022 6 · Michael Strobl, Amine Trabelsi, Osmar Zaïane

Modern approaches in Natural Language Processing (NLP) require, ideally, large amounts of labelled data for model training. However, new language resources, for example, for Named Entity Recognition (NER), Co-reference Resolution (CR), Entity Linking (EL) and Relation Extraction (RE), naming a few of the most popular tasks in NLP, have always been challenging to create since manual text annotations can be very time-consuming to acquire. While there may be an acceptable amount of labelled data available for some of these tasks in one language, there may be a lack of datasets in another. WEXEA is a tool to exhaustively annotate entities in the English Wikipedia. Guidelines for editors of Wikipedia articles result, on the one hand, in only a few annotations through hyperlinks, but on the other hand, make it easier to exhaustively annotate the rest of these articles with entities than starting from scratch. We propose the following main improvements to WEXEA: Creating multi-lingual corpora, improved entity annotations using a proven NER system, annotating dates and times. A brief evaluation of the annotation quality of WEXEA is added.

📄 PDF Abstract BibTeX

Code (1)

mjstrobl/wexea 공식 구현 tf

Tasks

ArticlesEntity Linkingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERRelation Extraction

Similar Papers 제목 키워드 기반

KnowNER: Incremental Multilingual Knowledge in Named Entity Recognition

2017-09-11 · Dominic Seyler, Tatiana Dembelova, Luciano del Corro, Johannes Hoffart 외

KnowNER is a multilingual Named Entity Recognition (NER) system that leverages different degrees of external knowledge. A novel modular framework divides the knowledge into four categories according to the depth of knowl…

Multilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Collaboratively Annotating Multilingual Parallel Corpora in the Biomedical Domain---some MANTRAs

2014-05-01 · LREC 2014 5 · Johannes Hellrich, Simon Clematide, Udo Hahn, Dietrich Rebholz-Schuhmann

The coverage of multilingual biomedical resources is high for the English language, yet sparse for non-English languages―an observation which holds for seemingly well-resourced, yet still dramatically low-resourced one…

Named Entity Recognition (NER)Translation

Fine-grained Named Entity Annotation for Finnish

2021-05-01 · NoDaLiDa 2021 5 · Jouni Luoma, Li-Hsin Chang, Filip Ginter, Sampo Pyysalo

We introduce a corpus with fine-grained named entity annotation for Finnish, following the OntoNotes guidelines to create a resource that is cross-lingually compatible with existing annotations for other languages. We co…

NER

MELO: An Evaluation Benchmark for Multilingual Entity Linking of Occupations

2024-10-10 · Federico Retyk, Luis Gasco, Casimiro Pio Carrino, Daniel Deniz 외

We present the Multilingual Entity Linking of Occupations (MELO) Benchmark, a new collection of 48 datasets for evaluating the linking of entity mentions in 21 languages to the ESCO Occupations multilingual taxonomy. MEL…

Entity LinkingSentence

FewTopNER: Integrating Few-Shot Learning with Topic Modeling and Named Entity Recognition in a Multilingual Framework

2025-02-04 · Ibrahim Bouabdallaoui, Fatima Guerouate, Samya Bouhaddour, Chaimae Saadi 외

We introduce FewTopNER, a novel framework that integrates few-shot named entity recognition (NER) with topic-aware contextual modeling to address the challenges of cross-lingual and low-resource scenarios. FewTopNER leve…

Entity DisambiguationFew-Shot Learningfew-shot-nerFew-shot NER+4