paper-with-me

Papers

Named Entity Corpus Construction using Wikipedia and DBpedia Ontology

2014-05-01 · LREC 2014 5 · Younggyun Hahm, Jungyeul Park, Kyungtae Lim, Youngsik Kim, Dosam Hwang, Key-Sun Choi

In this paper, we propose a novel method to automatically build a named entity corpus based on the DBpedia ontology. Since most of named entity recognition systems require time and effort consuming annotation tasks as training data. Work on NER has thus for been limited on certain languages like English that are resource-abundant in general. As an alternative, we suggest that the NE corpus generated by our proposed method, can be used as training data. Our approach introduces Wikipedia as a raw text and uses the DBpedia data set for named entity disambiguation. Our method is language-independent and easy to be applied to many different languages where Wikipedia and DBpedia are provided. Throughout the paper, we demonstrate that our NE corpus is of comparable quality even to the manually annotated NE corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Entity Disambiguationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

DBpedia Abstracts: A Large-Scale, Open, Multilingual NLP Training Corpus

2016-05-01 · LREC 2016 5 · Martin Br{\"u}mmer, Milan Dojchinovski, Sebastian Hellmann

The ever increasing importance of machine learning in Natural Language Processing is accompanied by an equally increasing need in large-scale training and evaluation corpora. Due to its size, its openness and relative qu…

Entity LinkingMultilingual NLP

Building and Evaluating Universal Named-Entity Recognition English corpus

2022-12-14 · Diego Alves, Gaurish Thakkar, Marko Tadić

This article presents the application of the Universal Named Entity framework to generate automatically annotated corpora. By using a workflow that extracts Wikipedia data and meta-data and DBpedia information, we genera…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Building Multilingual Corpora for a Complex Named Entity Recognition and Classification Hierarchy using Wikipedia and DBpedia

2022-12-14 · Diego Alves, Gaurish Thakkar, Gabriel Amaral, Tin Kuculo 외

With the ever-growing popularity of the field of NLP, the demand for datasets in low resourced-languages follows suit. Following a previously established framework, in this paper, we present the UNER dataset, a multiling…

ArticlesNamed Entity RecognitionNamed Entity Recognition (NER)

Building a Massive Corpus for Named Entity Recognition using Free Open Data Sources

2019-08-13 · Daniel Specht Menezes, Pedro Savarese, Ruy Luiz Milidiú

With the recent progress in machine learning, boosted by techniques such as deep learning, many tasks can be successfully solved once a large enough dataset is available for training. Nonetheless, human-annotated dataset…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

OPIEC: An Open Information Extraction Corpus

2019-04-28 · AKBC 2019 · Kiril Gashteovski, Sebastian Wanner, Sven Hertling, Samuel Broscheit 외

Open information extraction (OIE) systems extract relations and their arguments from natural language text in an unsupervised manner. The resulting extractions are a valuable resource for downstream tasks such as knowled…

Knowledge Base ConstructionOpen-Ended Question AnsweringOpen Information ExtractionQuestion Answering+1