Irregularity Detection in Categorized Document Corpora
The paper presents an approach to extract irregularities in document corpora, where the documents originate from different sources and the analyst's interest is to find documents which are atypical for the given source. The main contribution of the paper is a voting-based approach to irregularity detection and its evaluation on a collection of newspaper articles from two sources: Western (UK and US) and local (Kenyan) media. The evaluation of a domain expert proves that the method is very effective in uncovering interesting irregularities in categorized document corpora.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesDocument ClassificationOutlier DetectionText CategorizationSimilar Papers 제목 키워드 기반
AVID: Adversarial Visual Irregularity Detection
Real-time detection of irregularities in visual data is very invaluable and useful in many prospective applications including surveillance, patient monitoring systems, etc. With the surge of deep learning methods in the …
Anomaly DetectionCollaborative Imputation of Urban Time Series through Cross-city Meta-learning
Urban time series, such as mobility flows, energy consumption, and pollution records, encapsulate complex urban dynamics and structures. However, data collection in each city is impeded by technical challenges such as bu…
ImputationMeta-LearningTime SeriesImproving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling
Health conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at irregular time inte…
ImputationIrregular Time SeriesTime SeriesTime Series AnalysisIDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection
Effective fraud detection and analysis of government-issued identity documents, such as passports, driver's licenses, and identity cards, are essential in thwarting identity theft and bolstering security on online platfo…
Fraud DetectionPrivacy PreservingTime-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series
Time series data in real-world applications such as healthcare, climate modeling, and finance are often irregular, multimodal, and messy, with varying sampling rates, asynchronous modalities, and pervasive missingness. H…
Irregular Time SeriesTime SeriesTime Series Analysis