paper-with-me

홈 › Papers

Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

2023-09-29 · Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj

Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-training datasets and to mitigate its impact on downstream tasks. More specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) transfer performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing data distribution are different. We empirically verify that the reason behind is noise in pre-training shapes the feature space differently. We then propose a light-weight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering one may not be able to fully fine-tune or even access the pre-trained models. We conduct practical experiments on popular vision and language models that are pre-trained on noisy data for evaluation of our approach. Our analysis and results show the importance of this interesting and novel research direction, which we term Noisy Model Learning.

📄 PDF Abstract BibTeX arXiv:2309.17002

Code (1)

Hhhhhhao/Noisy-Model-Learning 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning

2024-03-10 · Zijun Long, Lipeng Zhuang, George Killick, Richard McCreadie 외

Human-annotated vision datasets inevitably contain a fraction of human mislabelled examples. While the detrimental effects of such mislabelling on supervised learning are well-researched, their influence on Supervised Co…

Contrastive LearningRepresentation Learning

Mitigating Noisy Inputs for Question Answering

2019-08-08 · Denis Peskov, Joe Barrow, Pedro Rodriguez, Graham Neubig 외

Natural language processing systems are often downstream of unreliable inputs: machine translation, optical character recognition, or speech recognition. For instance, virtual assistants can only answer your questions af…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationOptical Character Recognition+5

Skeleton-Based Human Action Recognition with Noisy Labels

2024-03-15 · Yi Xu, Kunyu Peng, Di Wen, Ruiping Liu 외

Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and a…

Action RecognitionDenoisingMixture-of-ExpertsSkeleton Based Action Recognition+2

Quantifying and mitigating the impact of label errors on model disparity metrics

2023-10-04 · Julius Adebayo, Melissa Hall, Bowen Yu, Bobbie Chern

Errors in labels obtained via human annotation adversely affect a model's performance. Existing approaches propose ways to mitigate the effect of label error on a model's downstream accuracy, yet little is known about it…

Bayesian Statistics Guided Label Refurbishment Mechanism: Mitigating Label Noise in Medical Image Classification

2021-06-23 · Mengdi Gao, Ximeng Feng, Mufeng Geng, Zhe Jiang 외

Purpose: Deep neural networks (DNNs) have been widely applied in medical image classification, benefiting from its powerful mapping capability among medical images. However, these existing deep learning-based methods dep…

Classificationimage-classificationImage ClassificationMedical Image Classification