A Publicly Available Annotated Corpus for Supervised Email Summarization
Annotated email corpora are necessary for evaluation and training of machine learning summarization techniques. The scarcity of corpora has been a limiting factor for research in this field. We describe our process of creating a new annotated email thread corpus that will be made publicly available. We present the trade-offs of the different annotation methods that could be used.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Crawling and Preprocessing Mailing Lists At Scale for Dialog Analysis
This paper introduces the Webis Gmane Email Corpus 2019, the largest publicly available and fully preprocessed email corpus to date. We crawled more than 153 million emails from 14,699 mailing lists and segmented them in…
Reply With: Proactive Recommendation of Email Attachments
Email responses often contain items-such as a file or a hyperlink to an external document-that are attached to or included inline in the body of the message. Analysis of an enterprise email corpus reveals that 35% of the…
Weakly-supervised LearningA Dataset for Anaphora Analysis in French Emails
In 2019, about 293 billion emails were sent worldwide every day. They are a valuable source of information and knowledge for professionals. Since the 90’s, many studies have been done on emails and have highlighted the n…
Building a Dataset for Summarization and Keyword Extraction from Emails
This paper introduces a new email dataset, consisting of both single and thread emails, manually annotated with summaries and keywords. A total of 349 emails and threads have been annotated. The dataset is our first step…
Abstractive Text SummarizationKeyword ExtractionA Study on Entity Resolution for Email Conversations
This paper investigates the problem of entity resolution for email conversations and presents a seed annotated corpus of email threads labeled with entity coreference chains. Characteristics of email threads concerning r…
Entity Resolution