Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
Pre-trained language models (PLMs) have achieved remarkable results on NLP tasks but at the expense of huge parameter sizes and the consequent computational costs. In this paper, we propose Variator, a parameter-efficient acceleration method that enhances computational efficiency through plug-and-play compression plugins. Compression plugins are designed to reduce the sequence length via compressing multiple hidden vectors into one and trained with original PLMs frozen. Different from traditional model acceleration methods, which compress PLMs to smaller sizes, Variator offers two distinct advantages: (1) In real-world applications, the plug-and-play nature of our compression plugins enables dynamic selection of different compression plugins with varying acceleration ratios based on the current workload. (2) The compression plugin comprises a few compact neural network layers with minimal parameters, significantly saving storage and memory overhead, particularly in scenarios with a growing number of tasks. We validate the effectiveness of Variator on seven datasets. Experimental results show that Variator can save 53% computational costs using only 0.9% additional parameters with a performance drop of less than 2%. Moreover, when the model scales to billions of parameters, Variator matches the strong performance of uncompressed PLMs.
Code (1)
Tasks
Computational EfficiencySimilar Papers 제목 키워드 기반
New class of compounds - variators - are reprogramming substrate specificity of H4K12Ac, H4K16Ac and H4K20Ac epigenetic marks reading bromodomain of BPTF protein
Previously reported [http://arxiv.org/abs/1506.06433] reprogramming of substrate specificity of H3K4Me3 epigenetic marks reading PHD domain of BPTF protein illustrates therapeutic potential of a new class of non-inhibito…
SpecificityDLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
In this paper, we propose the Dynamic Latent Frame Rate VAE (DLFR-VAE), a training-free paradigm that can make use of adaptive temporal compression in latent space. While existing video generative models apply fixed comp…
Video GenerationPeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator
We present Piecewise Rectified Flow (PeRFlow), a flow-based method for accelerating diffusion models. PeRFlow divides the sampling process of generative flows into several time windows and straightens the trajectories in…
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
Recent advancements in large language models (LLMs) are propelling us toward artificial general intelligence with their remarkable emergent abilities and reasoning capabilities. However, the substantial computational and…
BenchmarkingComputational EfficiencyLanguage ModelingLanguage Modelling+2New class of compounds - variators - are reprogramming substrate specificity of H3K4me3 epigenetic marks reading PHD domain of BPTF protein
In lymphoma, mutations in genes of histone modifying proteins are frequently observed. Notably, somatic mutations in the activatory histone modification writing protein MLL2 and the repressive modification writer EZH2 ar…
Specificity