paper-with-me

홈 › Papers

Content-based Models of Quotation

2021-04-01 · EACL 2021 2 · Ansel MacLaughlin, David Smith

We explore the task of quotability identification, in which, given a document, we aim to identify which of its passages are the most quotable, i.e. the most likely to be directly quoted by later derived documents. We approach quotability identification as a passage ranking problem and evaluate how well both feature-based and BERT-based (Devlin et al., 2019) models rank the passages in a given document by their predicted quotability. We explore this problem through evaluations on five datasets that span multiple languages (English, Latin) and genres of literature (e.g. poetry, plays, novels) and whose corresponding derived documents are of multiple types (news, journal articles). Our experiments confirm the relatively strong performance of BERT-based models on this task, with the best model, a RoBERTA sequential sentence tagger, achieving an average rho of 0.35 and NDCG@1, 5, 50 of 0.26, 0.31 and 0.40, respectively, across all five datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesPassage RankingSentence

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention 설명 없음
Adam 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Residual Connection 설명 없음
WordPiece 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Quotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existi…

CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction

2022-09-20 · COLING 2022 10 · Yequan Wang, Xiang Li, Aixin Sun, Xuying Meng 외

Quotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existi…

Continuity of Topic, Interaction, and Query: Learning to Quote in Online Conversations

2021-06-18 · EMNLP 2020 11 · Lingzhi Wang, Jing Li, Xingshan Zeng, Haisong Zhang 외

Quotations are crucial for successful explanations and persuasions in interpersonal communications. However, finding what to quote in a conversation is challenging for both humans and machines. This work studies automati…

DecoderText Generation

Extraction of unmarked quotations in Newspapers

2012-05-01 · LREC 2012 5 · St{\'e}phanie Weiser, Patrick Watrin

This paper presents work in progress to automatically extract quotation sentences from newspaper articles. The focus is the extraction and annotation of unmarked quotation sentences. A linguistic study shows that unmarke…

ArticlesOpinion MiningSentenceSpeech Recognition

The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts

2022-04-12 · LREC 2022 6 · Krishnapriya Vishnubhotla, Adam Hammond, Graeme Hirst

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of m…

Referring Expression