Crowdsourced Corpus with Entity Salience Annotations
In this paper, we present a crowdsourced dataset which adds entity salience (importance) annotations to the Reuters-128 dataset, which is subset of Reuters-21578. The dataset is distributed under a free license and publish in the NLP Interchange Format, which fosters interoperability and re-use. We show the potential of the dataset on the task of learning an entity salience classifier and report on the results from several experiments.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
WN-Salience: A Corpus of News Articles with Entity Salience Annotations
Entities can be found in various text genres, ranging from tweets and web pages to user queries submitted to web search engines. Existing research either considers all entities in the text equally important, or heuristic…
ArticlesEntity LinkingOpenPI2.0: An Improved Dataset for Entity Tracking in Texts
Much text describes a changing world (e.g., procedures, stories, newswires), and understanding them requires tracking how entities change. An earlier dataset, OpenPI, provided crowdsourced annotations of entity state cha…
Question AnsweringGUMsley: Evaluating Entity Salience in Summarization for 12 English Genres
As NLP models become increasingly capable of understanding documents in terms of coherent entities rather than strings, obtaining the most salient entities for each document is not only an important end task in itself bu…
Abstractive Text Summarizationcoreference-resolutionCoreference ResolutionHallucination+2Towards Better Text Understanding and Retrieval through Kernel Entity Salience Modeling
This paper presents a Kernel Entity Salience Model (KESM) that improves text understanding and retrieval by better estimating entity salience (importance) in documents. KESM represents entities by knowledge enriched dist…
RetrievalCrowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity Recognition
Crowdsourcing is regarded as one prospective solution for effective supervised learning, aiming to build large-scale annotated training data by crowd workers. Previous studies focus on reducing the influences from the no…
Domain Adaptationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2