paper-with-me

홈 › Papers

The Importance of Being Parameters: An Intra-Distillation Method for Serious Gains

2022-05-23 · Haoran Xu, Philipp Koehn, Kenton Murray

Recent model pruning methods have demonstrated the ability to remove redundant parameters without sacrificing model performance. Common methods remove redundant parameters according to the parameter sensitivity, a gradient-based measure reflecting the contribution of the parameters. In this paper, however, we argue that redundant parameters can be trained to make beneficial contributions. We first highlight the large sensitivity (contribution) gap among high-sensitivity and low-sensitivity parameters and show that the model generalization performance can be significantly improved after balancing the contribution of all parameters. Our goal is to balance the sensitivity of all parameters and encourage all of them to contribute equally. We propose a general task-agnostic method, namely intra-distillation, appended to the regular training loss to balance parameter sensitivity. Moreover, we also design a novel adaptive learning method to control the strength of intra-distillation loss for faster convergence. Our experiments show the strong effectiveness of our methods on machine translation, natural language understanding, and zero-shot cross-lingual transfer across up to 48 languages, e.g., a gain of 3.54 BLEU on average across 8 language pairs from the IWSLT'14 translation dataset.

📄 PDF Abstract BibTeX arXiv:2205.11416

Code (1)

fe1ixxu/intra-distillation 공식 구현 pytorch

Tasks

Cross-Lingual TransferMachine TranslationNatural Language UnderstandingSensitivityTranslationZero-Shot Cross-Lingual Transfer

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Language-Aware Multilingual Machine Translation with Self-Supervised Learning

2023-02-10 · Haoran Xu, Jean Maillard, Vedanuj Goswami

Multilingual machine translation (MMT) benefits from cross-lingual transfer but is a challenging multitask optimization problem. This is partly because there is no clear framework to systematically learn language-specifi…

Cross-Lingual TransferDecoderDenoisingMachine Translation+2

Importance-Aware Adaptive Dataset Distillation

2024-01-29 · Guang Li, Ren Togo, Takahiro Ogawa, Miki Haseyama

Herein, we propose a novel dataset distillation method for constructing small informative datasets that preserve the information of the large original datasets. The development of deep learning models is enabled by the a…

Dataset Distillation

Intra-class Patch Swap for Self-Distillation

2025-05-20 · Hongjun Choi, Eun Som Jeon, Ankita Shukla, Pavan Turaga

Knowledge distillation (KD) is a valuable technique for compressing large deep learning models into smaller, edge-suitable networks. However, conventional KD frameworks rely on pre-trained high-capacity teacher networks,…

image-classificationImage ClassificationKnowledge DistillationSemantic Segmentation

Explainable Rumor Detection using Inter and Intra-feature Attention Networks

2020-07-21 · Mingxuan Chen, Ning Wang, K. P. Subbalakshmi

With social media becoming ubiquitous, information consumption from this media has also increased. However, one of the serious problems that have emerged with this increase, is the propagation of rumors. Therefore, rumor…

Benchmarking

Distilling Model Knowledge

2015-10-08 · George Papamakarios

Top-performing machine learning systems, such as deep neural networks, large ensembles and complex probabilistic graphical models, can be expensive to store, slow to evaluate and hard to integrate into larger systems. Id…

Bayesian InferenceBIG-bench Machine LearningKnowledge Distillationmodel+1