TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection
Detecting novelty of an entire document is an Artificial Intelligence (AI) frontier problem that has widespread NLP applications, such as extractive document summarization, tracking development of news events, predicting impact of scholarly articles, etc. Important though the problem is, we are unaware of any benchmark document level data that correctly addresses the evaluation of automatic novelty detection techniques in a classification framework. To bridge this gap, we present here a resource for benchmarking the techniques for document level novelty detection. We create the resource via event-specific crawling of news documents across several domains in a periodic manner. We release the annotated corpus with necessary statistics and show its use with a developed system for the problem in concern.
Code (2)
Tasks
ArticlesBenchmarkingDocument SummarizationExtractive Document SummarizationExtractive Text SummarizationGeneral ClassificationNovelty DetectionSimilar Papers 제목 키워드 기반
NovAScore: A New Automated Metric for Evaluating Document Level Novelty
The rapid expansion of online content has intensified the issue of information redundancy, underscoring the need for solutions that can identify genuinely new information. Despite this challenge, the research community h…
Novelty DetectionNovelty Detection: A Perspective from Natural Language Processing
The quest for new information is an inborn human trait and has always been quintessential for human survival and progress. Novelty drives curiosity, which in turn drives innovation. In Natural Language Processing (NLP), …
Natural Language InferenceNovelty DetectionWord-level Human Interpretable Scoring Mechanism for Novel Text Detection Using Tsetlin Machines
Recent research in novelty detection focuses mainly on document-level classification, employing deep neural networks (DNN). However, the black-box nature of DNNs makes it difficult to extract an exact explanation of why …
Novelty DetectionText DetectionNovelty Goes Deep. A Deep Neural Solution To Document Level Novelty Detection
The rapid growth of documents across the web has necessitated finding means of discarding redundant documents and retaining novel ones. Capturing redundancy is challenging as it may involve investigating at a deep semant…
Document SummarizationFeature EngineeringInformation RetrievalNovelty Detection