paper-with-me

홈 › Papers

Revisiting Offline Compression: Going Beyond Factorization-based Methods for Transformer Language Models

2023-02-08 · Mohammadreza Banaei, Klaudia Bałazy, Artur Kasymov, Rémi Lebret, Jacek Tabor, Karl Aberer

Recent transformer language models achieve outstanding results in many natural language processing (NLP) tasks. However, their enormous size often makes them impractical on memory-constrained devices, requiring practitioners to compress them to smaller networks. In this paper, we explore offline compression methods, meaning computationally-cheap approaches that do not require further fine-tuning of the compressed model. We challenge the classical matrix factorization methods by proposing a novel, better-performing autoencoder-based framework. We perform a comprehensive ablation study of our approach, examining its different aspects over a diverse set of evaluation settings. Moreover, we show that enabling collaboration between modules across layers by compressing certain modules together positively impacts the final model performance. Experiments on various NLP tasks demonstrate that our approach significantly outperforms commonly used factorization-based offline compression methods.

📄 PDF Abstract BibTeX arXiv:2302.04045

Code (1)

mohammadrezabanaei/auto-encoder-based-transformer-compression 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs

2025-11-27 · Daniel Agyei Asante, Md Mokarram Chowdhury, Yang Li arxiv

Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained settings. Low-rank factorization addresses this challenge by compressing models to…

Adversarial Robustness

BALF: Budgeted Activation-Aware Low-Rank Factorization for Fine-Tuning-Free Model Compression

2025-09-29 · David González-Martínez arxiv

Activation-aware low-rank factorization techniques yield strong compression results but are generally confined to linear layers, while existing whitening-based theory typically makes an implicit full-rank assumption on a…

Model Compression

Compressed Nonnegative Matrix Factorization is Fast and Accurate

2015-05-18 · Mariano Tepper, Guillermo Sapiro

Nonnegative matrix factorization (NMF) has an established reputation as a useful data analysis technique in numerous applications. However, its usage in practical situations is undergoing challenges in recent years. The …

Lossless Model Compression via Joint Low-Rank Factorization Optimization

2024-12-09 · Boyang Zhang, Daning Cheng, Yunquan Zhang, Fangmin Liu 외

Low-rank factorization is a popular model compression technique that minimizes the error $\delta$ between approximated and original weight matrices. Despite achieving performances close to the original models when $\delt…

Model CompressionModel Optimization

Online Embedding Compression for Text Classification using Low Rank Matrix Factorization

2018-11-01 · Anish Acharya, Rahul Goel, Angeliki Metallinou, Inderjit Dhillon

Deep learning models have become state of the art for natural language processing (NLP) tasks, however deploying these models in production system poses significant memory constraints. Existing compression methods are ei…

General ClassificationQuantizationSentenceSentence Classification+2