Aleda, a free large-scale entity database for French
Named entity recognition, which focuses on the identification of the span and type of named entity mentions in texts, has drawn the attention of the NLP community for a long time. However, many real-life applications need to know which real entity each mention refers to. For such a purpose, often refered to as entity resolution and linking, an inventory of entities is required in order to constitute a reference. In this paper, we describe how we extracted such a resource for French from freely available resources (the French Wikipedia and the GeoNames database). We describe the results of an instrinsic evaluation of the resulting entity database, named Aleda, as well as those of a task-based evaluation in the context of a named entity detection system. We also compare it with the NLGbAse database (Charton and Torres-Moreno, 2010), a resource with similar objectives.
Code (0)
등록된 구현이 없습니다.
Tasks
Entity LinkingEntity ResolutionKnowledge Base Populationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Similar Papers 제목 키워드 기반
JRC-Names: A freely available, highly multilingual named entity resource
This paper describes a new, freely available, highly multilingual named entity resource for person and organisation names that has been compiled over seven years of large-scale multilingual news analysis combined with Wi…
Machine TranslationMorphological Inflectionnamed-entity-recognitionNamed Entity Recognition+2TypeNet: Deep Learning Keystroke Biometrics
We study the performance of Long Short-Term Memory networks for keystroke biometric authentication at large scale in free-text scenarios. For this we explore the performance of Long Short-Term Memory (LSTMs) networks tra…
Deep LearningTripletLarge-scale Taxonomy Induction Using Entity and Word Embeddings
Taxonomies are an important ingredient of knowledge organization, and serve as a backbone for more sophisticated knowledge representations in intelligent systems, such as formal ontologies. However, building taxonomies m…
Word Embeddingspreon: Fast and accurate entity normalization for drug names and cancer types in precision oncology
Motivation In precision oncology (PO), clinicians aim to find the best treatment for any patient based on their molecular characterization. A major bottleneck is the manual annotation and evaluation of individual varian…
Data IntegrationMedical Concept NormalizationTerm ExtractionEm-K Indexing for Approximate Query Matching in Large-scale ER
Accurate and efficient entity resolution (ER) is a significant challenge in many data mining and analysis projects requiring integrating and processing massive data collections. It is becoming increasingly important in r…
BlockingEntity Resolution