Hankel Singular Value Regularization for Highly Compressible State Space Models
Deep neural networks using state space models as layers are well suited for long-range sequence tasks but can be challenging to compress after training. We use that regularizing the sum of Hankel singular values of state space models leads to a fast decay of these singular values and thus to compressible models. To make the proposed Hankel singular value regularization scalable, we develop an algorithm to efficiently compute the Hankel singular values during training iterations by exploiting the specific block-diagonal structure of the system matrices that we use in our state space model parametrization. Experiments on Long Range Arena benchmarks demonstrate that the regularized state space layers are up to 10$\times$ more compressible than standard state space layers while maintaining high accuracy.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
System Identification via Nuclear Norm Regularization
This paper studies the problem of identifying low-order linear systems via Hankel nuclear norm regularization. Hankel regularization encourages the low-rankness of the Hankel matrix, which maps to the low-orderness of th…
Model SelectionFinite Sample System Identification: Improved Rates and the Role of Regularization
This paper studies low-order linear system identification via regularized regression. The nuclear norm of the system’s Hankel matrix is added as a regularizer to the least-squares cost function due to the following advan…
Hankel Singular Value Decomposition as a method of preprocessing the Magnetic Resonance Spectroscopy
The signal resulting from magnetic resonance spectroscopy is occupied by noises and irregularities so in the further analysis preprocessing techniques have to be introduced. The main idea of the paper is to develop a mod…
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
Deep neural networks achieve state of the art performance but remain difficult to interpret mechanistically. In this work, we propose a control theoretic framework that treats a trained neural network as a nonlinear stat…
On Low-Rank Hankel Matrix Denoising
The low-complexity assumption in linear systems can often be expressed as rank deficiency in data matrices with generalized Hankel structure. This makes it possible to denoise the data by estimating the underlying struct…
Denoising