paper-with-me

홈 › Papers

Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations

2022-10-14 · Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon, Jaewoo Kang

Most weakly supervised named entity recognition (NER) models rely on domain-specific dictionaries provided by experts. This approach is infeasible in many domains where dictionaries do not exist. While a phrase retrieval model was used to construct pseudo-dictionaries with entities retrieved from Wikipedia automatically in a recent study, these dictionaries often have limited coverage because the retriever is likely to retrieve popular entities rather than rare ones. In this study, we present a novel framework, HighGEN, that generates NER datasets with high-coverage pseudo-dictionaries. Specifically, we create entity-rich dictionaries with a novel search method, called phrase embedding search, which encourages the retriever to search a space densely populated with various entities. In addition, we use a new verification process based on the embedding distance between candidate entity mentions and entity types to reduce the false-positive noise in weak labels generated by high-coverage dictionaries. We demonstrate that HighGEN outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets.

📄 PDF Abstract BibTeX arXiv:2210.07586

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language QueriesNERRetrievalWeakly-Supervised Named Entity Recognition

Similar Papers 제목 키워드 기반

MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)

2022-07-01 · Findings (NAACL) 2022 7 · Simone Tedeschi, Roberto Navigli

Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datase…

Entity Linkingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

An Open Corpus for Named Entity Recognition in Historic Newspapers

2016-05-01 · LREC 2016 5 · Clemens Neudecker

The availability of openly available textual datasets ({``}corpora{''}) with highly accurate manual annotations ({``}gold standard{''}) of named entities (e.g. persons, locations, organizations, etc.) is crucial in the t…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

COVID-19 Named Entity Recognition for Vietnamese

2021-04-08 · NAACL 2021 4 · Thinh Hung Truong, Mai Hoang Dao, Dat Quoc Nguyen

The current COVID-19 pandemic has lead to the creation of many corpora that facilitate NLP research and downstream applications to help fight the pandemic. However, most of these corpora are exclusively for English. As t…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition In VietnameseNamed Entity Recognition (NER)+3

pioNER: Datasets and Baselines for Armenian Named Entity Recognition

2018-10-19 · Tsolak Ghukasyan, Garnik Davtyan, Karen Avetisyan, Ivan Andrianov

In this work, we tackle the problem of Armenian named entity recognition, providing silver- and gold-standard datasets as well as establishing baseline results on popular models. We present a 163000-token named entity co…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings

Automatic Creation of Arabic Named Entity Annotated Corpus Using Wikipedia

2014-04-01 · EACL 2014 4 · Maha Althobaiti, Udo Kruschwitz, Massimo Poesio
Morphological AnalysisNamed Entity Recognition (NER)