paper-with-me

홈 › Papers

AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression

2026-04-02 · Atul Kumar Sinha, François Fleuret arxiv

We introduce a fast low-rank factorization-based framework for compressing large language models that enables rapid compression of billion-parameter models without retraining. Unlike existing factorization-based approaches that optimize only on the original inputs, ignoring distribution shifts from upstream compression and thus propagating errors forward, or those that rely only on shifted inputs and risk drifting away from the original outputs, our approach accounts for both. Beyond individual layer compression, we further refine each transformer block end-to-end, minimizing block-level output distortion and allowing compressed layers to jointly compensate for accumulated errors. By anchoring each compressed layer to the original outputs while explicitly modeling input distribution shifts, our method finds a low-rank approximation that maintains functional equivalence with the original model. Experiments on large language models show that our method consistently outperforms existing SVD-based baselines across compression ratios, with the advantage becoming increasingly pronounced at aggressive compression budgets, where competing methods degrade substantially or collapse entirely, offering a practical solution for efficient, large-scale model deployment.

📄 PDF Abstract BibTeX arXiv:2604.02119

Code (0)

등록된 구현이 없습니다.

Tasks

Model Compression

Similar Papers 제목 키워드 기반

Sentence-Anchored Gist Compression for Long-Context LLMs

2025-11-11 · Dmitrii Tarasov, Elizaveta Goncharova, Kuznetsov Andrey arxiv

This work investigates context compression for Large Language Models (LLMs) using learned compression tokens to reduce the memory and computational demands of processing long sequences. We demonstrate that pre-trained LL…

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

2026-07-08 · Kyuan Oh, Bumsoo Kim arxiv

Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine qu…

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

2026-05-28 · Selim Kuzucu, Alessio Tonioni, Vasile Lup, Bernt Schiele 외 arxiv

Large Vision-Language Models (LVLMs) map visual inputs into dense token sequences, imposing a quadratic computational bottleneck for inference. Elastic visual-token compression addresses this by training a single model t…

KVSculpt: KV Cache Compression as Distillation

2026-03-29 · Bo Jiang, Sian Jin arxiv

KV cache compression is critical for efficient long-context LLM inference. Approaches that reduce the per-pair footprint -- quantization and low-rank decomposition -- are orthogonal to those that reduce the sequence leng…

Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model

2026-02-06 · Jacqueline He, Jonathan Hayase, Wen-tau Yih, Sewoong Oh 외 arxiv

Language models (LMs) tend to memorize portions of their training data and emit verbatim spans. When the underlying sources are sensitive or copyright-protected, such reproduction raises issues of consent and compensatio…