paper-with-me

홈 › Papers

Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It

2024-03-19 · Guoxuan Xia, Olivier Laurent, Gianni Franchi, Christos-Savvas Bouganis

Label smoothing (LS) is a popular regularisation method for training neural networks as it is effective in improving test accuracy and is simple to implement. Hard one-hot labels are smoothed by uniformly distributing probability mass to other classes, reducing overfitting. Prior work has suggested that in some cases LS can degrade selective classification (SC) -- where the aim is to reject misclassifications using a model's uncertainty. In this work, we first demonstrate empirically across an extended range of large-scale tasks and architectures that LS consistently degrades SC. We then address a gap in existing knowledge, providing an explanation for this behaviour by analysing logit-level gradients: LS degrades the uncertainty rank ordering of correct vs incorrect predictions by regularising the max logit more when a prediction is likely to be correct, and less when it is likely to be wrong. This elucidates previously reported experimental results where strong classifiers underperform in SC. We then demonstrate the empirical effectiveness of post-hoc logit normalisation for recovering lost SC performance caused by LS. Furthermore, linking back to our gradient analysis, we again provide an explanation for why such normalisation is effective.

📄 PDF Abstract BibTeX arXiv:2403.14715

Code (1)

ENSTA-U2IS-AI/Label-smoothing-Selective-classification-Code 공식 구현 pytorch

Tasks

Uncertainty Quantification

Similar Papers 제목 키워드 기반

Confidence-Aware Paced-Curriculum Learning by Label Smoothing for Surgical Scene Understanding

2022-12-22 · Mengya Xu, Mobarakol Islam, Ben Glocker, Hongliang Ren

Curriculum learning and self-paced learning are the training strategies that gradually feed the samples from easy to more complex. They have captivated increasing attention due to their excellent performance in robotic v…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONScene UnderstandingSemantic Segmentation

Pseudo Knowledge Distillation: Towards Learning Optimal Instance-specific Label Smoothing Regularization

2021-09-29 · Peng Lu, Ahmad Rashid, Ivan Kobyzev, Mehdi Rezagholizadeh 외

Knowledge Distillation (KD) is an algorithm that transfers the knowledge of a trained, typically larger, neural network into another model under training. Although a complete understanding of KD is elusive, a growing bod…

image-classificationImage ClassificationKnowledge DistillationNatural Language Understanding

Selective Output Smoothing Regularization: Regularize Neural Networks by Softening Output Distributions

2021-03-29 · Xuan Cheng, Tianshu Xie, Xiaomin Wang, Qifeng Weng 외

In this paper, we propose Selective Output Smoothing Regularization, a novel regularization method for training the Convolutional Neural Networks (CNNs). Inspired by the diverse effects on training from different samples…

image-classificationImage Classification

Boundary Smoothing for Named Entity Recognition

2022-04-26 · ACL 2022 5 · Enwei Zhu, Jinpeng Li

Neural named entity recognition (NER) models may easily encounter the over-confidence issue, which degrades the performance and calibration. Inspired by label smoothing and driven by the ambiguity of boundary annotation …

Chinese Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Boundary Smoothing for Named Entity Recognition

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Neural named entity recognition (NER) models may easily encounter the over-confidence issue, which degrades the performance and calibration. Inspired by label smoothing and driven by the ambiguity of boundary annotation …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER