paper-with-me

홈 › Papers

SMOClust: Synthetic Minority Oversampling based on Stream Clustering for Evolving Data Streams

2023-08-28 · Chun Wai Chiu, Leandro L. Minku

Many real-world data stream applications not only suffer from concept drift but also class imbalance. Yet, very few existing studies investigated this joint challenge. Data difficulty factors, which have been shown to be key challenges in class imbalanced data streams, are not taken into account by existing approaches when learning class imbalanced data streams. In this work, we propose a drift adaptable oversampling strategy to synthesise minority class examples based on stream clustering. The motivation is that stream clustering methods continuously update themselves to reflect the characteristics of the current underlying concept, including data difficulty factors. This nature can potentially be used to compress past information without caching data in the memory explicitly. Based on the compressed information, synthetic examples can be created within the region that recently generated new minority class examples. Experiments with artificial and real-world data streams show that the proposed approach can handle concept drift involving different minority class decomposition better than existing approaches, especially when the data stream is severely class imbalanced and presenting high proportions of safe and borderline minority class examples.

📄 PDF Abstract BibTeX arXiv:2308.14845

Code (1)

michaelchiucw/smoclust 공식 구현

Tasks

Clusteringimbalanced classificationSynthetic Data Generation

Similar Papers 제목 키워드 기반

Integrating Unsupervised Clustering and Label-specific Oversampling to Tackle Imbalanced Multi-label Data

2021-09-25 · Payel Sadhukhan, Arjun Pakrashi, Sarbani Palit, Brian Mac Namee

There is often a mixture of very frequent labels and very infrequent labels in multi-label datatsets. This variation in label frequency, a type class imbalance, creates a significant challenge for building efficient mult…

ClusteringMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

GMOTE: Gaussian based minority oversampling technique for imbalanced classification adapting tail probability of outliers

2021-05-09 · Seung Jee Yang, Kyung Joon Cha

Classification of imbalanced data is one of the common problems in the recent field of data mining. Imbalanced data substantially affects the performance of standard classification models. Data-level approaches mainly us…

ClusteringGeneral Classificationimbalanced classification

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

2025-10-23 · Touqeer Ahmad, Mohammadreza M. Kalan, François Portier, Gilles Stupfler arxiv

Synthetic oversampling of minority examples using SMOTE and its variants is a leading strategy for addressing imbalanced classification problems. Despite the success of this approach in practice, its theoretical foundati…

A Novel Adaptive Minority Oversampling Technique for Improved Classification in Data Imbalanced Scenarios

2021-03-24 · Ayush Tripathi, Rupayan Chakraborty, Sunil Kumar Kopparapu

Imbalance in the proportion of training samples belonging to different classes often poses performance degradation of conventional classifiers. This is primarily due to the tendency of the classifier to be biased towards…

ClusteringGeneral Classification

A multi-schematic classifier-independent oversampling approach for imbalanced datasets

2021-07-15 · Saptarshi Bej, Kristian Schultz, Prashant Srivastava, Markus Wolfien 외

Over 85 oversampling algorithms, mostly extensions of the SMOTE algorithm, have been built over the past two decades, to solve the problem of imbalanced datasets. However, it has been evident from previous studies that d…

Benchmarking