paper-with-me

홈 › Papers

Identifying Hard Noise in Long-Tailed Sample Distribution

2022-07-27 · Xuanyu Yi, Kaihua Tang, Xian-Sheng Hua, Joo-Hwee Lim, Hanwang Zhang

Conventional de-noising methods rely on the assumption that all samples are independent and identically distributed, so the resultant classifier, though disturbed by noise, can still easily identify the noises as the outliers of training distribution. However, the assumption is unrealistic in large-scale data that is inevitably long-tailed. Such imbalanced training data makes a classifier less discriminative for the tail classes, whose previously "easy" noises are now turned into "hard" ones -- they are almost as outliers as the clean tail samples. We introduce this new challenge as Noisy Long-Tailed Classification (NLT). Not surprisingly, we find that most de-noising methods fail to identify the hard noises, resulting in significant performance drop on the three proposed NLT benchmarks: ImageNet-NLT, Animal10-NLT, and Food101-NLT. To this end, we design an iterative noisy learning framework called Hard-to-Easy (H2E). Our bootstrapping philosophy is to first learn a classifier as noise identifier invariant to the class and context distributional changes, reducing "hard" noises to "easy" ones, whose removal further improves the invariance. Experimental results show that our H2E outperforms state-of-the-art de-noising methods and their ablations on long-tailed settings while maintaining a stable performance on the conventional balanced settings. Datasets and codes are available at https://github.com/yxymessi/H2E-Framework

📄 PDF Abstract BibTeX arXiv:2207.13378

Code (1)

yxymessi/h2e-framework 공식 구현 pytorch

Tasks

Philosophy

Similar Papers 제목 키워드 기반

Label-Noise Learning with Intrinsically Long-Tailed Data

2022-08-21 · ICCV 2023 1 · Yang Lu, Yiliang Zhang, Bo Han, Yiu-ming Cheung 외

Label noise is one of the key factors that lead to the poor generalization of deep learning models. Existing label-noise learning methods usually assume that the ground-truth classes of the training data are balanced. Ho…

Disentangling Hardness from Noise: An Uncertainty-Driven Model-Agnostic Framework for Long-Tailed Remote Sensing Classification

2026-01-01 · Chi Ding, Junxiao Xue, Xinyi Yin, Shi Chen 외 arxiv

Long-Tailed distributions are pervasive in remote sensing due to the inherently imbalanced occurrence of grounded objects. However, a critical challenge remains largely overlooked, i.e., disentangling hard tail data samp…

Sample hardness based gradient loss for long-tailed cervical cell detection

2022-08-07 · Minmin Liu, Xuechen Li, Xiangbo Gao, Junliang Chen 외

Due to the difficulty of cancer samples collection and annotation, cervical cancer datasets usually exhibit a long-tailed data distribution. When training a detector to detect the cancer cells in a WSI (Whole Slice Image…

Cell Detectionobject-detectionObject Detection

Co-Learning Meets Stitch-Up for Noisy Multi-label Visual Recognition

2023-07-03 · Chao Liang, Zongxin Yang, Linchao Zhu, Yi Yang

In real-world scenarios, collected and annotated data often exhibit the characteristics of multiple classes and long-tailed distribution. Additionally, label noise is inevitable in large-scale annotations and hinders the…

Learning with noisy labelsMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONRepresentation Learning

TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly Detection

2025-04-03 · CVPR 2025 1 · Yoon Gyo Jung, Jaewoo Park, Jaeho Yoon, Kuan-Chuan Peng 외

We aim to solve unsupervised anomaly detection in a practical challenging environment where the normal dataset is both contaminated with defective regions and its product class distribution is tailed but unknown. We obse…

Anomaly DetectionUnsupervised Anomaly Detection