paper-with-me

Papers

Centroids: Gold standards with distributional variation

2012-05-01 · LREC 2012 5 · Ian Lewin, {\c{S}}enay Kafkas, Dietrich Rebholz-Schuhmann

Motivation: Gold Standards for named entities are, ironically, not standard themselves. Some specify the “one perfect annotation”. Others specify “perfectly good alternatives”. The concept of Silver standard is relatively new. The objective is consensus rather than perfection. How should the two concepts be best represented and related? Approach: We examine several Biomedical Gold Standards and motivate a new representational format, centroids, which simply and effectively represents name distributions. We define an algorithm for finding centroids, given a set of alternative input annotations and we test the outputs quantitatively and qualitatively. We also define a metric of relatively acceptability on top of the centroid standard. Results: Precision, recall and F-scores of over 0.99 are achieved for the simple sanity check of giving the algorithm Gold Standard inputs. Qualitative analysis of the differences very often reveals errors and incompleteness in the original Gold Standard. Given automatically generated annotations, the centroids effectively represent the range of those contributions and the quality of the centroid annotations is highly competitive with the best of the contributors. Conclusion: Centroids cleanly represent alternative name variations for Silver and Gold Standards. A centroid Silver Standard is derived just like a Gold Standard, only from imperfect inputs.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Named Entity Recognition (NER)

Similar Papers 제목 키워드 기반

SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation

2014-08-15 · CL 2015 12 · Felix Hill, Roi Reichart, Anna Korhonen

We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and ME…

Representation Learning

Exposing ambiguities in a relation-extraction gold standard with crowdsourcing

2015-05-23 · Tong Shu Li, Benjamin M. Good, Andrew I. Su

Semantic relation extraction is one of the frontiers of biomedical natural language processing research. Gold standards are key tools for advancing this research. It is challenging to generate these standards because of …

RelationRelation Extraction

Gollum: A Gold Standard for Large Scale Multi Source Knowledge Graph Matching

2022-09-15 · Sven Hertling, Heiko Paulheim

The number of Knowledge Graphs (KGs) generated with automatic and manual approaches is constantly growing. For an integrated view and usage, an alignment between these KGs is necessary on the schema as well as instance l…

Graph MatchingKnowledge Graphs

Neural Text Classification by Jointly Learning to Cluster and Align

2020-11-24 · Yekun Chai, Haidong Zhang, Shuo Jin

Distributional text clustering delivers semantically informative representations and captures the relevance between each word and semantic clustering centroids. We extend the neural text clustering approach to text class…

ClassificationClusteringGeneral Classificationtext-classification+3

Analysing Inconsistencies and Errors in PoS Tagging in two Icelandic Gold Standards

2015-05-01 · WS 2015 5 · Stein{\th}{\'o}r Steingr{\'\i}msson, Sigr{\'u}n Helgad{\'o}ttir, Eir{\'\i}kur R{\"o}gnvaldsson
POSPOS Tagging