paper-with-me

홈 › Papers

PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient Training

2026-02-26 · Yanyi Li, Yimu Zhang, Cong Fang arxiv

Activations have become the primary memory bottleneck in large-batch LLM training. However, existing compression methods fail to exploit the spectral structure of activations, resulting in slow convergence or limited compression. To address this, we bridge the relationship between the algorithm's fast convergence and the requirements for subspace projection, and show that an effective compression should yield an unbiased estimate of the original activation with low variance. We propose Principal-Random Subspace for LLM Activation Compression (PRAC), which novelly decomposes activations into two components: a principal subspace captured via SVD to retain dominant information, and a random subspace sampled from the orthogonal complement to approximate the tail. By introducing a precise scaling factor, we prove that PRAC yields an unbiased gradient estimator with minimum variance under certain conditions. Extensive experiments on pre-training and fine-tuning tasks demonstrate that PRAC achieves up to 36% total memory reduction with negligible performance degradation and minimal computational cost.

📄 PDF Abstract BibTeX arXiv:2602.23111

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Novel Stochastic Gradient Descent Algorithm for Learning Principal Subspaces

2022-12-08 · Charline Le Lan, Joshua Greaves, Jesse Farebrother, Mark Rowland 외

Many machine learning problems encode their data as a matrix with a possibly very large number of rows and columns. In several applications like neuroscience, image compression or deep reinforcement learning, the princip…

Deep Reinforcement LearningImage Compressionreinforcement-learningReinforcement Learning (RL)

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

2024-12-18 · Utkarsh Saxena, Sayeh Sharify, Kaushik Roy, Xin Wang

Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors…

Quantization

LASER: Low-Rank Activation SVD for Efficient Recursion

2026-04-19 · Ege Çakar, Ketan Ali Raghu, Lia Zheng arxiv

Recursive architectures such as Tiny Recursive Models (TRMs) perform implicit reasoning through iterative latent computation, yet the geometric structure of these reasoning trajectories remains poorly understood. We inve…

LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning

2025-09-25 · Marco Paul E. Apolinario, Kaushik Roy arxiv

On-device learning is essential for personalization, privacy, and long-term adaptation in resource-constrained environments. Achieving this requires efficient learning, both fine-tuning existing models and continually ac…

Continual Learning

From Principal Subspaces to Principal Components with Linear Autoencoders

2018-04-26 · Elad Plaut

The autoencoder is an effective unsupervised learning model which is widely used in deep learning. It is well known that an autoencoder with a single fully-connected hidden layer, a linear activation function and a squar…

Dimensionality Reduction