paper-with-me

홈 › Papers

Compression Method for Deep Diagonal State Space Model Based on $H^2$ Optimal Reduction

2025-07-14 · Hiroki Sakamoto, Kazuhiro Sato arxiv

Deep learning models incorporating linear SSMs have gained attention for capturing long-range dependencies in sequential data. However, their large parameter sizes pose challenges for deployment on resource-constrained devices. In this study, we propose an efficient parameter reduction method for these models by applying $H^{2}$ model order reduction techniques from control theory to their linear SSM components. In experiments, the LRA benchmark results show that the model compression based on our proposed method outperforms an existing method using the Balanced Truncation, while successfully reducing the number of parameters in the SSMs to $1/32$ without sacrificing the performance of the original models.

📄 PDF Abstract BibTeX arXiv:2507.10078

Code (0)

등록된 구현이 없습니다.

Tasks

Model Compression

Similar Papers 제목 키워드 기반

Model Compression Method for S4 with Diagonal State Space Layers using Balanced Truncation

2024-02-25 · Haruka Ezoe, Kazuhiro Sato

To implement deep learning models on edge devices, model compression methods have been widely recognized as useful. However, it remains unclear which model compression methods are effective for Structured State Space Seq…

Model Compression

Optimal Diagonal Preconditioning

2022-09-02 · Zhaonan Qu, Wenzhi Gao, Oliver Hinder, Yinyu Ye 외

Preconditioning has long been a staple technique in optimization, often applied to reduce the condition number of a matrix and speed up the convergence of algorithms. Although there are many popular preconditioning techn…

SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices

2026-06-05 · Ernests Lavrinovics, Marco Letizia, Roy Janco, Shai Segal 외 arxiv

We present SigmaScale, a method for learning auxiliary scaling matrices $S$ to aid truncated Singular Value Decomposition (SVD) based Large Language Model (LLM) compression. Instead of deriving scaling matrices analytica…

ESPACE: Dimensionality Reduction of Activations for Model Compression

2024-10-07 · Charbel Sakr, Brucek Khailany

We propose ESPACE, an LLM compression technique based on dimensionality reduction of activations. Unlike prior works on weight-centric tensor decomposition, ESPACE projects activations onto a pre-calibrated set of princi…

Dimensionality ReductionmodelModel CompressionTensor Decomposition

Artemis: HE-Aware Training for Efficient Privacy-Preserving Machine Learning

2023-10-02 · Yeonsoo Jeon, Mattan Erez, Michael Orshansky

Privacy-Preserving ML (PPML) based on Homomorphic Encryption (HE) is a promising foundational privacy technology. Making it more practical requires lowering its computational cost, especially, in handling modern large de…

Model CompressionPrivacy Preserving