Content-based Models of Quotation
We explore the task of quotability identification, in which, given a document, we aim to identify which of its passages are the most quotable, i.e. the most likely to be directly quoted by later derived documents. We approach quotability identification as a passage ranking problem and evaluate how well both feature-based and BERT-based (Devlin et al., 2019) models rank the passages in a given document by their predicted quotability. We explore this problem through evaluations on five datasets that span multiple languages (English, Latin) and genres of literature (e.g. poetry, plays, novels) and whose corresponding derived documents are of multiple types (news, journal articles). Our experiments confirm the relatively strong performance of BERT-based models on this task, with the best model, a RoBERTA sequential sentence tagger, achieving an average rho of 0.35 and NDCG@1, 5, 50 of 0.26, 0.31 and 0.40, respectively, across all five datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesPassage RankingSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction
Quotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existi…
CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction
Quotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existi…
Continuity of Topic, Interaction, and Query: Learning to Quote in Online Conversations
Quotations are crucial for successful explanations and persuasions in interpersonal communications. However, finding what to quote in a conversation is challenging for both humans and machines. This work studies automati…
DecoderText GenerationExtraction of unmarked quotations in Newspapers
This paper presents work in progress to automatically extract quotation sentences from newspaper articles. The focus is the extraction and annotation of unmarked quotation sentences. A linguistic study shows that unmarke…
ArticlesOpinion MiningSentenceSpeech RecognitionThe Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts
We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of m…
Referring Expression