paper-with-me

홈 › Papers

Frobenius-Type Norms and Inner Products of Matrices and Linear Maps with Applications to Neural Network Training

2023-11-26 · Roland Herzog, Frederik Köhne, Leonie Kreis, Anton Schiela

The Frobenius norm is a frequent choice of norm for matrices. In particular, the underlying Frobenius inner product is typically used to evaluate the gradient of an objective with respect to matrix variable, such as those occuring in the training of neural networks. We provide a broader view on the Frobenius norm and inner product for linear maps or matrices, and establish their dependence on inner products in the domain and co-domain spaces. This shows that the classical Frobenius norm is merely one special element of a family of more general Frobenius-type norms. The significant extra freedom furnished by this realization can be used, among other things, to precondition neural network training.

📄 PDF Abstract BibTeX arXiv:2311.15419

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

2026-05-31 · Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra arxiv

The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling is challenging due to the presence of additional symmetries under co…

Consensus Seminorms and their Applications

2025-05-07 · Ron Ofir, Ji Liu, A. Stephen Morse, Brian D. O. Anderson

Consensus is a well-studied problem in distributed sensing, computation and control, yet deriving useful and easily computable bounds on the rate of convergence to consensus remains a challenge. We study the applications…

Minimax Estimation of Quadratic Fourier Functionals

2018-03-30 · Shashank Singh, Bharath K. Sriperumbudur, Barnabás Póczos

We study estimation of (semi-)inner products between two nonparametric probability distributions, given IID samples from each distribution. These products include relatively well-studied classical $\mathcal{L}^2$ and Sob…

Translation

More for Keys, Less for Values: Adaptive KV Cache Quantization

2025-02-20 · Mohsen Hariri, Lam Nguyen, Sixu Chen, Shaochen Zhong 외

This paper introduces an information-aware quantization framework that adaptively compresses the key-value (KV) cache in large language models (LLMs). Although prior work has underscored the distinct roles of key and val…

Quantization

On Size-Independent Sample Complexity of ReLU Networks

2023-06-03 · Mark Sellke

We study the sample complexity of learning ReLU neural networks from the point of view of generalization. Given norm constraints on the weight matrices, a common approach is to estimate the Rademacher complexity of the a…