paper-with-me

홈 › Papers

RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation

2021-09-21 · Findings (NAACL) 2022 7 · Md Akmal Haidar, Nithin Anchuri, Mehdi Rezagholizadeh, Abbas Ghaddar, Philippe Langlais, Pascal Poupart

Intermediate layer knowledge distillation (KD) can improve the standard KD technique (which only targets the output of teacher and student models) especially over large pre-trained language models. However, intermediate layer distillation suffers from excessive computational burdens and engineering efforts required for setting up a proper layer mapping. To address these problems, we propose a RAndom Intermediate Layer Knowledge Distillation (RAIL-KD) approach in which, intermediate layers from the teacher model are selected randomly to be distilled into the intermediate layers of the student model. This randomized selection enforce that: all teacher layers are taken into account in the training process, while reducing the computational cost of intermediate layer distillation. Also, we show that it act as a regularizer for improving the generalizability of the student model. We perform extensive experiments on GLUE tasks as well as on out-of-domain test sets. We show that our proposed RAIL-KD approach outperforms other state-of-the-art intermediate layer KD methods considerably in both performance and training-time.

📄 PDF Abstract BibTeX arXiv:2109.10164

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Methods 이 논문이 사용한 방법론

Test 설명 없음
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover's Distance

2020-10-13 · EMNLP 2020 11 · Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu 외

Pre-trained language models (e.g., BERT) have achieved significant success in various natural language processing (NLP) tasks. However, high storage and computational costs obstruct pre-trained language models to be effe…

Model Compression

Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge Distillation

2021-11-01 · EMNLP 2021 11 · Yimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md Akmal Haidar 외

Intermediate layer matching is shown as an effective approach for improving knowledge distillation (KD). However, this technique applies matching in the hidden spaces of two different networks (i.e. student and teacher),…

Knowledge Distillation

CoT Rerailer: Enhancing the Reliability of Large Language Models in Complex Reasoning Tasks through Error Detection and Correction

2024-08-25 · Guangya Wan, Yuqi Wu, Jie Chen, Sheng Li

Chain-of-Thought (CoT) prompting enhances Large Language Models (LLMs) complex reasoning abilities by generating intermediate steps. However, these steps can introduce hallucinations and accumulate errors. We propose the…

Decision MakingQuestion Answering

How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

2024-06-09 · Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu 외

Large language models (LLMs) rely on safety alignment to avoid responding to malicious user inputs. Unfortunately, jailbreak can circumvent safety guardrails, resulting in LLMs generating harmful content and raising conc…

Safety Alignment

Conversion of Braille to Text in English, Hindi and Tamil Languages

2013-07-11 · S. Padmavathi, Manojna K. S. S, S. Sphoorthy Reddy, D. Meenakshy

The Braille system has been used by the visually impaired for reading and writing. Due to limited availability of the Braille text books an efficient usage of the books becomes a necessity. This paper proposes a method t…