paper-with-me

홈 › Papers

SeismoFlow -- Data augmentation for the class imbalance problem

2020-07-23 · Ruy Luiz Milidiú, Luis Felipe Müller

In several application areas, such as medical diagnosis, spam filtering, fraud detection, and seismic data analysis, it is very usual to find relevant classification tasks where some class occurrences are rare. This is the so called class imbalance problem, which is a challenge in machine learning. In this work, we propose the SeismoFlow a flow-based generative model to create synthetic samples, aiming to address the class imbalance. Inspired by the Glow model, it uses interpolation on the learned latent space to produce synthetic samples for one rare class. We apply our approach to the development of a seismogram signal quality classifier. We introduce a dataset composed of5.223seismograms that are distributed between the good, medium, and bad classes and with their respective frequencies of 66.68%,31.54%, and 1.76%. Our methodology is evaluated on a stratified 10-fold cross-validation setting, using the Miniceptionmodel as a baseline, and assessing the effects of adding the generated samples on the training set of each iteration. In our experiments, we achieve an improvement of 13.9% on the rare class F1-score, while not hurting the metric value for the other classes and thus observing the overall accuracy improvement. Our empirical findings indicate that our method can generate high-quality synthetic seismograms with realistic looking and sufficient plurality to help the Miniception model to overcome the class imbalance problem. We believe that our results are a step forward in solving both the task of seismogram signal quality classification and class imbalance.

📄 PDF Abstract BibTeX arXiv:2007.12229

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationFraud DetectionMedical Diagnosis

Methods 이 논문이 사용한 방법론

Normalizing Flows Normalizing Flows are a method for constructing complex distributions by transforming a probability density through a series of invertible mappings. By repeatedly applying…
Affine Coupling 설명 없음
Activation Normalization Activation Normalization is a type of normalization used for flow-based generative models; specifically it was introduced in the GLOW…
Invertible 1x1 Convolution The Invertible 1x1 Convolution is a type of convolution used in flow-based generative models that reverses the ordering of…
GLOW 설명 없음

Similar Papers 제목 키워드 기반

A review of ensemble learning and data augmentation models for class imbalanced problems: combination, implementation and evaluation

2023-04-06 · Azal Ahmad Khan, Omkar Chaudhari, Rohitash Chandra

Class imbalance (CI) in classification problems arises when the number of observations belonging to one class is lower than the other. Ensemble learning combines multiple models to obtain a robust model and has been prom…

Data AugmentationEnsemble Learning

Improving Model Performance and Removing the Class Imbalance Problem Using Augmentation

2022-05-01 · International Journal of Advanced Research in Engineering and Technology (IJARET) 2022 5 · Allena Venkata Sai Abhishek, Dr. Venkateswara Rao Gurrala

The data in the real world consists of various kinds of painful features. A majorly found one is the class imbalance in which the number of examples in different classes in a dataset is unequal. The class imbalance is be…

ClassificationData AugmentationData VisualizationDetecting Image Manipulation+13

CAISA at SemEval-2023 Task 8: Counterfactual Data Augmentation for Mitigating Class Imbalance in Causal Claim Identification

2023-06-01 · Akbar Karimi, Lucie Flek

The class imbalance problem can cause machine learning models to produce an undesirable performance on the minority class as well as the whole dataset. Using data augmentation techniques to increase the number of samples…

counterfactualData Augmentation

CUDA: Curriculum of Data Augmentation for Long-Tailed Recognition

2023-02-10 · Sumyeong Ahn, Jongwoo Ko, Se-Young Yun

Class imbalance problems frequently occur in real-world tasks, and conventional deep learning algorithms are well known for performance degradation on imbalanced training datasets. To mitigate this problem, many approach…

Data AugmentationLong-tail Learning

Empirical Study of Text Augmentation on Social Media Text in Vietnamese

2020-09-25 · Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

In the text classification problem, the imbalance of labels in datasets affect the performance of the text-classification models. Practically, the data about user comments on social networking sites not altogether appear…

Data AugmentationGeneral ClassificationHate Speech DetectionSentiment Analysis+4