paper-with-me

홈 › Papers

Phrase Detectives Corpus 1.0 Crowdsourced Anaphoric Coreference.

2016-05-01 · LREC 2016 5 · Jon Chamberlain, Massimo Poesio, Udo Kruschwitz

Natural Language Engineering tasks require large and complex annotated datasets to build more advanced models of language. Corpora are typically annotated by several experts to create a gold standard; however, there are now compelling reasons to use a non-expert crowd to annotate text, driven by cost, speed and scalability. Phrase Detectives Corpus 1.0 is an anaphorically-annotated corpus of encyclopedic and narrative text that contains a gold standard created by multiple experts, as well as a set of annotations created by a large non-expert crowd. Analysis shows very good inter-expert agreement (kappa=.88-.93) but a more variable baseline crowd agreement (kappa=.52-.96). Encyclopedic texts show less agreement (and by implication are harder to annotate) than narrative texts. The release of this corpus is intended to encourage research into the use of crowds for text annotation and the development of more advanced, probabilistic language models, in particular for anaphoric coreference.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

text annotation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

A Crowdsourced Corpus of Multiple Judgments and Disagreement on Anaphoric Interpretation

2019-06-01 · NAACL 2019 6 · Massimo Poesio, Jon Chamberlain, Silviu Paun, Juntao Yu 외

We present a corpus of anaphoric information (coreference) crowdsourced through a game-with-a-purpose. The corpus, containing annotations for about 108,000 markables, is one of the largest corpora for coreference for Eng…

Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts

2022-10-11 · Juntao Yu, Silviu Paun, Maris Camilleri, Paloma Carretero Garcia 외

Although several datasets annotated for anaphoric reference/coreference exist, even the largest such datasets have limitations in terms of size, range of domains, coverage of anaphoric phenomena, and size of documents in…

2k

A Probabilistic Annotation Model for Crowdsourcing Coreference

2018-10-01 · EMNLP 2018 10 · Silviu Paun, Jon Chamberlain, Udo Kruschwitz, Juntao Yu 외

The availability of large scale annotated corpora for coreference is essential to the development of the field. However, creating resources at the required scale via expert annotation would be too expensive. Crowdsourcin…

Coreference ResolutionmodelQuestion Answering

Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations

2026-06-24 · Anna Nedoluzhko, Šárka Zikánová, Jiří Mírovský, Milan Straka 외 arxiv

As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to another. To explore this phenomenon, we created two corpora with mu…

Coreference Resolution

Towards Coreference Resolution for Early Irish

2022-06-01 · CLTW (LREC) 2022 6 · Mark Darling, Marieke Meelen, David Willis

In this article, we present an outline of some of the issues involved in developing a semi-supervised procedure for coreference resolution for early Irish as part of a wider enterprise to create a parsed corpus of histor…

coreference-resolutionCoreference Resolution