paper-with-me

홈 › Papers

IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression

2026-05-15 · Ali Abbasi, Chayne Thrash, Haoran Qin, Hamed Pirsiavash, Soheil Kolouri arxiv

Large language models deliver strong performance across language and reasoning tasks, but their storage and compute costs remain major barriers to deployment in resource-constrained and latency-sensitive settings. SVD-based post-training compression offers a hardware-agnostic way to reduce model size and improve inference efficiency through low-rank factorization. However, existing methods often rely on input-only whitening spaces, homogeneous rank allocation, or loss-agnostic allocation heuristics, limiting their ability to preserve model quality under aggressive compression. We propose Input-Output Whitened SVD (IO-SVD), a post-training compression method that forms a KL-aware double-sided whitening space for model weights. Using a second-order expansion of the KL loss over the top-K token probabilities, IO-SVD constructs an output-side metric that captures predictive sensitivity, while input whitening captures activation statistics. We further introduce an efficient heterogeneous rank-allocation strategy that scores whitened singular components using first-order calibration loss estimates and prunes the least sensitive components under a global budget. Inspired by prior work that combines SVD truncation with quantization, we improve hybrid SVD-quantization compression through loss-aware remapping, which selects low-rank factor rows for 8-bit quantization based on the predicted loss change incurred by quantizing them. Extensive experiments across diverse LLM and VLM families, and inference-time analysis shows that IO-SVD compresses LLMs with minimal performance degradation while delivering practical inference speedups. Code is available at https://github.com/mint-vu/IO-SVD.git

📄 PDF Abstract BibTeX arXiv:2605.15626

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression

2026-04-02 · Atul Kumar Sinha, François Fleuret arxiv

We introduce a fast low-rank factorization-based framework for compressing large language models that enables rapid compression of billion-parameter models without retraining. Unlike existing factorization-based approach…

Model Compression

Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression

2026-02-02 · Ali Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma 외 arxiv

Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storage and can speed up inference via low-ran…

MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs

2025-07-04 · Guangyan Li, Yongqiang Tang, Wensheng Zhang arxiv

The enormous parameter scale of large language models (LLMs) has made model compression a research hotspot, which aims to alleviate computational resource demands during deployment and inference. As a promising direction…

Model Compression

SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression

2026-02-03 · Xing Hu, Dawei Yang, Yuan Cheng, Zhixuan Chen 외 arxiv

The rapid growth in the parameter scale of large language models (LLMs) has created a high demand for efficient compression techniques. As a hardware-agnostic and highly compatible technique, low-rank compression has bee…

Optimal Brain Decomposition for Accurate LLM Low-Rank Approximation

2026-04-01 · Yuhang Li, Donghyun Lee, Ruokai Yin, Priyadarshini Panda arxiv

Low-rank decomposition has emerged as an important problem in Large Language Model (LLM) fine-tuning and inference. Through Singular Value Decomposition (SVD), the weight matrix can be factorized into low-rank spaces opt…