paper-with-me

Papers

DirectQuote: A Dataset for Direct Quotation Extraction and Attribution in News Articles

2021-10-15 · LREC 2022 6 · Yuanchi Zhang, Yang Liu

Quotation extraction and attribution are challenging tasks, aiming at determining the spans containing quotations and attributing each quotation to the original speaker. Applying this task to news data is highly related to fact-checking, media monitoring and news tracking. Direct quotations are more traceable and informative, and therefore of great significance among different types of quotations. Therefore, this paper introduces DirectQuote, a corpus containing 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media. To the best of our knowledge, this is the largest and most complete corpus that focuses on direct quotations in news texts. We ensure that each speaker in the annotation can be linked to a specific named entity on Wikidata, benefiting various downstream tasks. In addition, for the first time, we propose several sequence labeling models as baseline methods to extract and attribute quotations simultaneously in an end-to-end manner.

📄 PDF Abstract BibTeX arXiv:2110.07827

Code (1)

thunlp-mt/directquote 공식 구현

Tasks

ArticlesAttributeFact Checking

Similar Papers 제목 키워드 기반

FRACAS: A FRench Annotated Corpus of Attribution relations in newS

2023-09-19 · Ange Richard, Laura Alonzo-Canul, François Portet

Quotation extraction is a widely useful task both from a sociological and from a Natural Language Processing perspective. However, very little data is available to study this task in languages other than English. In this…

Identifying Speakers and Addressees of Quotations in Novels with Prompt Learning

2024-08-18 · Yuchen Yan, Hanjie Zhao, Senbin Zhu, Hongde Liu 외

Quotations in literary works, especially novels, are important to create characters, reflect character relationships, and drive plot development. Current research on quotation extraction in novels primarily focuses on qu…

Prompt Learning

The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts

2022-04-12 · LREC 2022 6 · Krishnapriya Vishnubhotla, Adam Hammond, Graeme Hirst

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of m…

Referring Expression

Improving Automatic Quotation Attribution in Literary Novels

2023-07-07 · Krishnapriya Vishnubhotla, Frank Rudzicz, Graeme Hirst, Adam Hammond

Current models for quotation attribution in literary novels assume varying levels of available information in their training and test data, which poses a challenge for in-the-wild inference. Here, we approach quotation a…

coreference-resolutionCoreference Resolution

Dataset of Quotation Attribution in German News Articles

2024-04-25 · Fynn Petersen-Frey, Chris Biemann

Extracting who says what to whom is a crucial part in analyzing human communication in today's abundance of data such as online news articles. Yet, the lack of annotated data for this task in German news articles severel…

Articles