Data Separability for Neural Network Classifiers and the Development of a Separability Index
In machine learning, the performance of a classifier depends on both the classifier model and the dataset. For a specific neural network classifier, the training process varies with the training set used; some training data make training accuracy fast converged to high values, while some data may lead to slowly converged to lower accuracy. To quantify this phenomenon, we created the Distance-based Separability Index (DSI), which is independent of the classifier model, to measure the separability of datasets. In this paper, we consider the situation where different classes of data are mixed together in the same distribution is most difficult for classifiers to separate, and we show that the DSI can indicate whether data belonging to different classes have similar distributions. When comparing our proposed approach with several existing separability/complexity measures using synthetic and real datasets, the results show the DSI is an effective separability measure. We also discussed possible applications of the DSI in the fields of data science, machine learning, and deep learning.
Code (1)
Tasks
BIG-bench Machine LearningSimilar Papers 제목 키워드 기반
A Novel Intrinsic Measure of Data Separability
In machine learning, the performance of a classifier depends on both the classifier model and the separability/complexity of datasets. To quantitatively measure the separability of datasets, we create an intrinsic measur…
The Role of Subgroup Separability in Group-Fair Medical Image Classification
We investigate performance disparities in deep classifiers. We find that the ability of classifiers to separate individuals into subgroups varies substantially across medical imaging modalities and protected characterist…
image-classificationImage ClassificationMedical Image ClassificationDCSI -- An improved measure of cluster separability based on separation and connectedness
Whether class labels in a given data set correspond to meaningful clusters is crucial for the evaluation of clustering algorithms using real-world data sets. This property can be quantified by separability measures. The …
ClusteringAn Internal Cluster Validity Index Using a Distance-based Separability Measure
To evaluate clustering results is a significant part of cluster analysis. There are no true class labels for clustering in typical unsupervised learning. Thus, a number of internal evaluations, which use predicted labels…
ClusteringClustering Algorithms EvaluationLarge-Scale Quantum Separability Through a Reproducible Machine Learning Lens
The quantum separability problem consists in deciding whether a bipartite density matrix is entangled or separable. In this work, we propose a machine learning pipeline for finding approximate solutions for this NP-hard …
Benchmarking