paper-with-me

홈 › Papers

JRC-Names: A freely available, highly multilingual named entity resource

2013-09-24 · Ralf Steinberger, Bruno Pouliquen, Mijail Kabadjov, Erik van der Goot

This paper describes a new, freely available, highly multilingual named entity resource for person and organisation names that has been compiled over seven years of large-scale multilingual news analysis combined with Wikipedia mining, resulting in 205,000 per-son and organisation names plus about the same number of spelling variants written in over 20 different scripts and in many more languages. This resource, produced as part of the Europe Media Monitor activity (EMM, http://emm.newsbrief.eu/overview.html), can be used for a number of purposes. These include improving name search in databases or on the internet, seeding machine learning systems to learn named entity recognition rules, improve machine translation results, and more. We describe here how this resource was created; we give statistics on its current size; we address the issue of morphological inflection; and we give details regarding its functionality. Updates to this resource will be made available daily.

📄 PDF Abstract BibTeX arXiv:1309.6162

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationMorphological Inflectionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Translation

Similar Papers 제목 키워드 기반

ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata

2024-05-15 · Jonne Sälevä, Constantine Lignos

We introduce ParaNames, a massively multilingual parallel name resource consisting of 140 million names spanning over 400 languages. Names are provided for 16.8 million entities, and each entity is mapped from a complex …

Multilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionTranslation+1

ParaNames: A Massively Multilingual Entity Name Corpus

2022-02-28 · NAACL (SIGTYP) 2022 7 · Jonne Sälevä, Constantine Lignos

We introduce ParaNames, a multilingual parallel name resource consisting of 118 million names spanning across 400 languages. Names are provided for 13.6 million entities which are mapped to standardized entity types (PER…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Translation+1

PhoBERT: Pre-trained language models for Vietnamese

2020-03-02 · Findings of the Association for Computational Linguistics 2020 · Dat Quoc Nguyen, Anh Tuan Nguyen

We present PhoBERT with two versions, PhoBERT-base and PhoBERT-large, the first public large-scale monolingual language models pre-trained for Vietnamese. Experimental results show that PhoBERT consistently outperforms t…

Dependency Parsingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages

2024-12-12 · Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 외

We present OpenNER 1.0, a standardized collection of openly available named entity recognition (NER) datasets. OpenNER contains 34 datasets spanning 51 languages, annotated in varying named entity ontologies. We correct …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

MSNER: A Multilingual Speech Dataset for Named Entity Recognition

2024-05-19 · Quentin Meeus, Marie-Francine Moens, Hugo Van hamme

While extensively explored in text-based tasks, Named Entity Recognition (NER) remains largely neglected in spoken language understanding. Existing resources are limited to a single, English-only dataset. This paper addr…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1