paper-with-me

홈 › Papers

Intriguing Properties of Compression on Multilingual Models

2022-11-04 · Kelechi Ogueji, Orevaoghene Ahia, Gbemileke Onilude, Sebastian Gehrmann, Sara Hooker, Julia Kreutzer

Multilingual models are often particularly dependent on scaling to generalize to a growing number of languages. Compression techniques are widely relied upon to reconcile the growth in model size with real world resource constraints, but compression can have a disparate effect on model performance for low-resource languages. It is thus crucial to understand the trade-offs between scale, multilingualism, and compression. In this work, we propose an experimental framework to characterize the impact of sparsifying multilingual pre-trained language models during fine-tuning. Applying this framework to mBERT named entity recognition models across 40 languages, we find that compression confers several intriguing and previously unknown generalization properties. In contrast to prior findings, we find that compression may improve model robustness over dense models. We additionally observe that under certain sparsification regimes compression may aid, rather than disproportionately impact the performance of low-resource languages.

📄 PDF Abstract BibTeX arXiv:2211.02738

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Methods 이 논문이 사용한 방법론

mBERT mBERT

Similar Papers 제목 키워드 기반

IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?

2024-10-03 · Akhilesh Aravapalli, Mounika Marreddy, Subba Reddy Oota, Radhika Mamidi 외

Transformer-based models have revolutionized the field of natural language processing. To understand why they perform so well and to assess their reliability, several studies have focused on questions such as: Which ling…

Privacy, Interpretability, and Fairness in the Multilingual Space

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Multilingual generalization or compression is an objective for cross-lingual models in natural language processing (NLP). We explore how the compression sought for in such models aligns with other common objectives in NL…

FairnessRetrievalSentenceSentence Retrieval

RuSentEval: Linguistic Source, Encoder Force!

2021-02-28 · EACL (BSNLP) 2021 4 · Vladislav Mikhailov, Ekaterina Taktasheva, Elina Sigdel, Ekaterina Artemova

The success of pre-trained transformer language models has brought a great deal of interest on how these models work, and what they learn about language. However, prior research in the field is mainly devoted to English,…

Sample Compression, Support Vectors, and Generalization in Deep Learning

2018-11-05 · Christopher Snyder, Sriram Vishwanath

Even though Deep Neural Networks (DNNs) are widely celebrated for their practical performance, they possess many intriguing properties related to depth that are difficult to explain both theoretically and intuitively. Un…

Deep Learning

Multilingual Brain Surgeon: Large Language Models Can be Compressed Leaving No Language Behind

2024-04-06 · Hongchuan Zeng, Hongshen Xu, Lu Chen, Kai Yu

Large Language Models (LLMs) have ushered in a new era in Natural Language Processing, but their massive size demands effective compression techniques for practicality. Although numerous model compression techniques have…

Model Compression