How to detect novelty in textual data streams? A comparative study of existing methods
Since datasets with annotation for novelty at the document and/or word level are not easily available, we present a simulation framework that allows us to create different textual datasets in which we control the way novelty occurs. We also present a benchmark of existing methods for novelty detection in textual data streams. We define a few tasks to solve and compare several state-of-the-art methods. The simulation framework allows us to evaluate their performances according to a set of limited scenarios and test their sensitivity to some parameters. Finally, we experiment with the same methods on different kinds of novelty in the New York Times Annotated Dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Novelty DetectionSensitivitySimilar Papers 제목 키워드 기반
An Improved System for Sentence-level Novelty Detection in Textual Streams
Novelty detection in news events has long been a difficult problem. A number of models performed well on specific data streams but certain issues are far from being solved, particularly in large data streams from the WWW…
Event DetectionNovelty DetectionSentenceParameterizing Kterm Hashing
Kterm Hashing provides an innovative approach to novelty detection on massive data streams. Previous research focused on maximizing the efficiency of Kterm Hashing and succeeded in scaling First Story Detection to Twitte…
Novelty DetectionWhere's Wally Now? Deep Generative and Discriminative Embeddings for Novelty Detection
We develop a framework for novelty detection (ND) methods relying on deep embeddings, either discriminative or generative, and also propose a novel framework for assessing their performance. While much progress was made…
Novelty DetectionA comparative evaluation of novelty detection algorithms for discrete sequences
The identification of anomalies in temporal data is a core component of numerous research areas such as intrusion detection, fault prevention, genomics and fraud detection. This article provides an experimental compariso…
Fraud DetectionIntrusion DetectionNovelty Detection