paper-with-me

홈 › Papers

CALBC: Releasing the Final Corpora

2012-05-01 · LREC 2012 5 · {\c{S}}enay Kafkas, Ian Lewin, David Milward, Erik van Mulligen, Jan Kors, Udo Hahn, Dietrich Rebholz-Schuhmann

A number of gold standard corpora for named entity recognition are available to the public. However, the existing gold standard corpora are limited in size and semantic entity types. These usually lead to implementation of trained solutions (1) for a limited number of semantic entity types and (2) lacking in generalization capability. In order to overcome these problems, the CALBC project has aimed to automatically generate large scale corpora annotated with multiple semantic entity types in a community-wide manner based on the consensus of different named entity solutions. The generated corpus is called the silver standard corpus since the corpus generation process does not involve any manual curation. In this publication, we announce the release of the final CALBC corpora which include the silver standard corpus in different versions and several gold standard corpora for the further usage of the biomedical text mining community. The gold standard corpora are utilised to benchmark the methods used in the silver standard corpora generation process and released in a shared format. All the corpora are released in a shared format and accessible at www.calbc.eu.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

HBert + BiasCorp -- Fighting Racism on the Web

2021-04-06 · Olawale Onabola, Zhuang Ma, Yang Xie, Benjamin Akera 외

Subtle and overt racism is still present both in physical and online communities today and has impacted many lives in different segments of the society. In this short piece of work, we present how we're tackling this soc…

hBERT + BiasCorp - Fighting Racism on the Web

2021-04-01 · EACL (LTEDI) 2021 4 · Olawale Onabola, Zhuang Ma, Xie Yang, Benjamin Akera 외

Subtle and overt racism is still present both in physical and online communities today and has impacted many lives in different segments of the society. In this short piece of work, we present how we’re tackling this soc…

Differentially Private Releasing via Deep Generative Model (Technical Report)

2018-01-05 · Xinyang Zhang, Shouling Ji, Ting Wang

Privacy-preserving releasing of complex data (e.g., image, text, audio) represents a long-standing challenge for the data mining research community. Due to rich semantics of the data and lack of a priori knowledge about …

Privacy Preserving

Robust Detection of Synthetic Tabular Data under Schema Variability

2025-08-27 · G. Charbel N. Kindji, Elisa Fromont, Lina Maria Rojas-Barahona, Tanguy Urvoy arxiv

The rise of powerful generative models has sparked concerns over data authenticity. While detection methods have been extensively developed for images and text, the case of tabular data, despite its ubiquity, has been la…

Misspelling Oblivious Word Embeddings

2019-05-23 · NAACL 2019 6 · Bora Edizel, Aleksandra Piktus, Piotr Bojanowski, Rui Ferreira 외

In this paper we present a method to learn word embeddings that are resilient to misspellings. Existing word embeddings have limited applicability to malformed texts, which contain a non-negligible amount of out-of-vocab…

Word Embeddings