paper-with-me

홈 › Papers

SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model Compression

2026-03-12 · Xinhao Huang, You-Liang Huang, Zeyi Wen arxiv

Large language models (LLMs) have demonstrated impressive capabilities across various tasks, but the billion-scale parameters pose deployment challenges. Although existing methods attempt to reduce the scale of LLMs, they require either special hardware support or expensive post-training to maintain model quality. To facilitate efficient and affordable model slimming, we propose a novel training-free compression method for LLMs, named "SoLA", which leverages \textbf{So}ft activation sparsity and \textbf{L}ow-r\textbf{A}nk decomposition. SoLA can identify and retain a minority of components significantly contributing to inference, while compressing the majority through low-rank decomposition, based on our analysis of the activation pattern in the feed-forward network (FFN) of modern LLMs. To alleviate the decomposition loss, SoLA is equipped with an adaptive component-wise low-rank allocation strategy to assign appropriate truncation positions for different weight matrices. We conduct extensive experiments on LLaMA-2-7B/13B/70B and Mistral-7B models across a variety of benchmarks. SoLA exhibits remarkable improvement in both language modeling and downstream task accuracy without post-training. For example, with a 30\% compression rate on the LLaMA-2-70B model, SoLA surpasses the state-of-the-art method by reducing perplexity from 6.95 to 4.44 and enhancing downstream task accuracy by 10\%.

📄 PDF Abstract BibTeX arXiv:2604.03258

Code (0)

등록된 구현이 없습니다.

Tasks

Model Compression

Similar Papers 제목 키워드 기반

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity

2026-05-05 · Jiaxi Li, Lu Yin, Li Shen, Jinjin Xu 외 arxiv

Large Language Models (LLMs) have achieved remarkable capabilities, but their immense computational demands during training remain a critical bottleneck for widespread adoption. Low-rank training has received attention i…

Deep Neural Network Initialization with Sparsity Inducing Activations

2024-02-25 · Ilan Price, Nicholas Daultry Ball, Samuel C. H. Lam, Adam C. Jones 외

Inducing and leveraging sparse activations during training and inference is a promising avenue for improving the computational efficiency of deep networks, which is increasingly important as network sizes continue to gro…

Computational Efficiency

Sparsifying Networks via Subdifferential Inclusion

2021-01-01 · Sagar Verma, Jean-Christophe Pesquet

Sparsifying deep neural networks is of paramount interest in many areas, especially when those networks have to be implemented on low-memory devices. In this article, we propose a new formulation of the problem of genera…

image-classificationImage Classificationspeech-recognitionSpeech Recognition+4

Softpick: No Attention Sink, No Massive Activations with Rectified Softmax

2025-04-29 · Zayd M. K. Zuhri, Erland Hilman Fuadi, Alham Fikri Aji

We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention sink and massive activations. Our experiments with 340M parameter models d…

Quantization

Solar: $L_0$ solution path averaging for fast and accurate variable selection in high-dimensional data

2020-07-30 · Ning Xu, Timothy C. G. Fisher

We propose a new variable selection algorithm, subsample-ordered least-angle regression (solar), and its coordinate descent generalization, solar-cd. Solar re-constructs lasso paths using the $L_0$ norm and averages the …

Variable Selection