paper-with-me

홈 › Papers

Knowledge Distillation with Noisy Labels for Natural Language Understanding

2021-09-21 · WNUT (ACL) 2021 11 · Shivendra Bhardwaj, Abbas Ghaddar, Ahmad Rashid, Khalil Bibi, Chengyang Li, Ali Ghodsi, Philippe Langlais, Mehdi Rezagholizadeh

Knowledge Distillation (KD) is extensively used to compress and deploy large pre-trained language models on edge devices for real-world applications. However, one neglected area of research is the impact of noisy (corrupted) labels on KD. We present, to the best of our knowledge, the first study on KD with noisy labels in Natural Language Understanding (NLU). We document the scope of the problem and present two methods to mitigate the impact of label noise. Experiments on the GLUE benchmark show that our methods are effective even under high noise levels. Nevertheless, our results indicate that more research is necessary to cope with label noise under the KD.

📄 PDF Abstract BibTeX arXiv:2109.10147

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationNatural Language Understanding

Similar Papers 제목 키워드 기반

FlyKD: Graph Knowledge Distillation on the Fly with Curriculum Learning

2024-03-16 · Eugene Ku

Knowledge Distillation (KD) aims to transfer a more capable teacher model's knowledge to a lighter student model in order to improve the efficiency of the model, making it faster and more deployable. However, the student…

Knowledge Distillation

Blind Knowledge Distillation for Robust Image Classification

2022-11-21 · Timo Kaiser, Lukas Ehmann, Christoph Reinders, Bodo Rosenhahn

Optimizing neural networks with noisy labels is a challenging task, especially if the label set contains real-world noise. Networks tend to generalize to reasonable patterns in the early training stages and overfit to sp…

Classificationimage-classificationImage ClassificationKnowledge Distillation+1

Federated Learning with Extremely Noisy Clients via Negative Distillation

2023-12-20 · Yang Lu, Lin Chen, Yonggang Zhang, Yiliang Zhang 외

Federated learning (FL) has shown remarkable success in cooperatively training deep models, while typically struggling with noisy labels. Advanced works propose to tackle label noise by a re-weighting strategy with a str…

Federated LearningKnowledge Distillation

Student as an Inherent Denoiser of Noisy Teacher

2023-12-15 · Jiachen Zhao

Knowledge distillation (KD) has been widely employed to transfer knowledge from a large language model (LLM) to a specialized model in low-data regimes through pseudo label learning. However, pseudo labels generated by t…

Knowledge DistillationLanguage ModelingLanguage ModellingLarge Language Model+1

Learning from Noisy Labels with Distillation

2017-03-07 · ICCV 2017 10 · Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao 외

The ability of learning from noisy labels is very useful in many visual recognition tasks, as a vast amount of data with noisy labels are relatively easy to obtain. Traditionally, the label noises have been treated as st…