paper-with-me

홈 › Papers

Assessing the Quality of the Datasets by Identifying Mislabeled Samples

2021-09-10 · Vaibhav Pulastya, Gaurav Nuti, Yash Kumar Atri, Tanmoy Chakraborty

Due to the over-emphasize of the quantity of data, the data quality has often been overlooked. However, not all training data points contribute equally to learning. In particular, if mislabeled, it might actively damage the performance of the model and the ability to generalize out of distribution, as the model might end up learning spurious artifacts present in the dataset. This problem gets compounded by the prevalence of heavily parameterized and complex deep neural networks, which can, with their high capacity, end up memorizing the noise present in the dataset. This paper proposes a novel statistic -- noise score, as a measure for the quality of each data point to identify such mislabeled samples based on the variations in the latent space representation. In our work, we use the representations derived by the inference network of data quality supervised variational autoencoder (AQUAVS). Our method leverages the fact that samples belonging to the same class will have similar latent representations. Therefore, by identifying the outliers in the latent space, we can find the mislabeled samples. We validate our proposed statistic through experimentation by corrupting MNIST, FashionMNIST, and CIFAR10/100 datasets in different noise settings for the task of identifying mislabelled samples. We further show significant improvements in accuracy for the classification task for each dataset.

📄 PDF Abstract BibTeX arXiv:2109.05000

Code (1)

lcs2-iiitd/aquavs 공식 구현 tf

Similar Papers 제목 키워드 기반

Identifying Mislabeled Data using the Area Under the Margin Ranking

2020-01-28 · NeurIPS 2020 12 · Geoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. Weinberger

Not all data in a typical training set help with generalization; some samples can be overly ambiguous or outrightly mislabeled. This paper introduces a new method to identify such samples and mitigate their impact when t…

Learning from Training Dynamics: Identifying Mislabeled Data Beyond Manually Designed Features

2022-12-19 · Qingrui Jia, Xuhong LI, Lei Yu, Jiang Bian 외

While mislabeled or ambiguously-labeled samples in the training set could negatively affect the performance of deep models, diagnosing the dataset and identifying mislabeled samples helps to improve the generalization po…

Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models

2024-05-06 · Anshuman Chhabra, Bo Li, Jian Chen, Prasant Mohapatra 외

A core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for ass…

On Revisiting Entropy for Identifying Mislabeled Images

2026-05-29 · Chunlei Li, Zixuan Zheng, Yilei Shi, Guanglu Dong 외 arxiv

Mislabeled samples in training datasets severely degrade the performance of deep networks, as overparameterized models tend to memorize erroneous labels. We address this challenge by proposing a novel approach for mislab…

Computational Efficiency

Identifying the Mislabeled Training Samples of ECG Signals using Machine Learning

2017-12-11 · Yaoguang Li, Wei Cui, Cong Wang

The classification accuracy of electrocardiogram signal is often affected by diverse factors in which mislabeled training samples issue is one of the most influential problems. In order to mitigate this negative effect, …

BIG-bench Machine LearningClassificationGeneral Classification