paper-with-me

홈 › Papers

A Study imbalance handling by various data sampling methods in binary classification

2021-05-23 · Mohamed Hamama

The purpose of this research report is to present the our learning curve and the exposure to the Machine Learning life cycle, with the use of a Kaggle binary classification data set and taking to explore various techniques from pre-processing to the final optimization and model evaluation, also we highlight on the data imbalance issue and we discuss the different methods of handling that imbalance on the data level by over-sampling and under sampling not only to reach a balanced class representation but to improve the overall performance. This work also opens some gaps for future work.

📄 PDF Abstract BibTeX arXiv:2105.10959

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningBinary Classification

Similar Papers 제목 키워드 기반

Handling Imbalanced Data: A Case Study for Binary Class Problems

2020-10-09 · Richmond Addo Danquah

For several years till date, the major issues in terms of solving for classification problems are the issues of Imbalanced data. Because majority of the machine learning algorithms by default assumes all data are balance…

Binary Classification

Balancing the Scales: A Comprehensive Study on Tackling Class Imbalance in Binary Classification

2024-09-29 · Mohamed Abdelhamid, Abhyuday Desai

Class imbalance in binary classification tasks remains a significant challenge in machine learning, often resulting in poor performance on minority classes. This study comprehensively evaluates three widely-used strategi…

Binary Classification

Rebalancing the Scales: A Systematic Mapping Study of Generative Adversarial Networks (GANs) in Addressing Data Imbalance

2025-02-23 · Pankaj Yadav, Gulshan Sihag, Vivek Vijay

Machine learning algorithms are used in diverse domains, many of which face significant challenges due to data imbalance. Studies have explored various approaches to address the issue, like data preprocessing, cost-sensi…

Survey of Imbalanced Data Methodologies

2021-04-06 · Lian Yu, Nengfeng Zhou

Imbalanced data set is a problem often found and well-studied in financial industry. In this paper, we reviewed and compared some popular methodologies handling data imbalance. We then applied the under-sampling/over-sam…

Survey

CSMOUTE: Combined Synthetic Oversampling and Undersampling Technique for Imbalanced Data Classification

2020-04-07 · Michał Koziarski

In this paper we propose a novel data-level algorithm for handling data imbalance in the classification task, Synthetic Majority Undersampling Technique (SMUTE). SMUTE leverages the concept of interpolation of nearby ins…

General Classification