SeismoFlow -- Data augmentation for the class imbalance problem
In several application areas, such as medical diagnosis, spam filtering, fraud detection, and seismic data analysis, it is very usual to find relevant classification tasks where some class occurrences are rare. This is the so called class imbalance problem, which is a challenge in machine learning. In this work, we propose the SeismoFlow a flow-based generative model to create synthetic samples, aiming to address the class imbalance. Inspired by the Glow model, it uses interpolation on the learned latent space to produce synthetic samples for one rare class. We apply our approach to the development of a seismogram signal quality classifier. We introduce a dataset composed of5.223seismograms that are distributed between the good, medium, and bad classes and with their respective frequencies of 66.68%,31.54%, and 1.76%. Our methodology is evaluated on a stratified 10-fold cross-validation setting, using the Miniceptionmodel as a baseline, and assessing the effects of adding the generated samples on the training set of each iteration. In our experiments, we achieve an improvement of 13.9% on the rare class F1-score, while not hurting the metric value for the other classes and thus observing the overall accuracy improvement. Our empirical findings indicate that our method can generate high-quality synthetic seismograms with realistic looking and sufficient plurality to help the Miniception model to overcome the class imbalance problem. We believe that our results are a step forward in solving both the task of seismogram signal quality classification and class imbalance.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationFraud DetectionMedical DiagnosisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A review of ensemble learning and data augmentation models for class imbalanced problems: combination, implementation and evaluation
Class imbalance (CI) in classification problems arises when the number of observations belonging to one class is lower than the other. Ensemble learning combines multiple models to obtain a robust model and has been prom…
Data AugmentationEnsemble LearningImproving Model Performance and Removing the Class Imbalance Problem Using Augmentation
The data in the real world consists of various kinds of painful features. A majorly found one is the class imbalance in which the number of examples in different classes in a dataset is unequal. The class imbalance is be…
ClassificationData AugmentationData VisualizationDetecting Image Manipulation+13CAISA at SemEval-2023 Task 8: Counterfactual Data Augmentation for Mitigating Class Imbalance in Causal Claim Identification
The class imbalance problem can cause machine learning models to produce an undesirable performance on the minority class as well as the whole dataset. Using data augmentation techniques to increase the number of samples…
counterfactualData AugmentationCUDA: Curriculum of Data Augmentation for Long-Tailed Recognition
Class imbalance problems frequently occur in real-world tasks, and conventional deep learning algorithms are well known for performance degradation on imbalanced training datasets. To mitigate this problem, many approach…
Data AugmentationLong-tail LearningEmpirical Study of Text Augmentation on Social Media Text in Vietnamese
In the text classification problem, the imbalance of labels in datasets affect the performance of the text-classification models. Practically, the data about user comments on social networking sites not altogether appear…
Data AugmentationGeneral ClassificationHate Speech DetectionSentiment Analysis+4