paper-with-me

Papers

Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization

2025-12-10 · Alexey Kravatskiy, Ivan Kozyrev, Nikolai Kozlov, Alexander Vinogradov, Daniil Merkulov, Ivan Oseledets arxiv

In this article, we explore the use of various matrix norms for optimizing functions of weight matrices, a crucial problem in deep learning. Moving beyond the spectral norm that underlies the Muon update, we leverage the duals of the Ky Fan norms to introduce the Fanion family of linear minimization oracle (LMO) algorithms, which are closely related to Muon, $ν$-SAM, and Dion. Staying inside the LMO, we construct the families of F-Fanions and S-Fanions, whose updates are convex combinations of the updates of Fanions and Normalized SGD or SignSGD, respectively. The most promising algorithms in these families are F-Muon and S-Muon. By conducting an extensive empirical study of all three algorithm families across a wide range of tasks and settings, we demonstrate that F-Muon and S-Muon consistently match Muon's performance, while outperforming Muon on a synthetic smooth convex problem.

📄 PDF Abstract BibTeX arXiv:2512.09678

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Matrix reconstruction with the local max norm

2012-12-01 · NeurIPS 2012 12 · Rina Foygel, Nathan Srebro, Ruslan R. Salakhutdinov

We introduce a new family of matrix norms, the ''local max'' norms, generalizing existing methods such as the max norm, the trace norm (nuclear norm), and the weighted or smoothed weighted trace norms, which have been ex…

Implicit Regularization in Deep Learning May Not Be Explainable by Norms

2020-05-13 · NeurIPS 2020 12 · Noam Razin, Nadav Cohen

Mathematically characterizing the implicit regularization induced by gradient-based optimization is a longstanding pursuit in the theory of deep learning. A widespread hope is that a characterization based on minimizatio…

Deep LearningMatrix CompletionOpen-Ended Question Answering

Convex Coupled Matrix and Tensor Completion

2017-05-15 · Kishan Wimalawarne, Makoto Yamada, Hiroshi Mamitsuka

We propose a set of convex low rank inducing norms for a coupled matrices and tensors (hereafter coupled tensors), which shares information between matrices and tensors through common modes. More specifically, we propose…

The Singular Value Decomposition, Applications and Beyond

2015-10-29 · Zhihua Zhang

The singular value decomposition (SVD) is not only a classical theory in matrix computation and analysis, but also is a powerful tool in machine learning and modern data analysis. In this tutorial we first study the basi…

BIG-bench Machine LearningMatrix Completion

Spectral k-Support Norm Regularization

2014-12-01 · NeurIPS 2014 12 · Andrew M. McDonald, Massimiliano Pontil, Dimitris Stamos

The $k$-support norm has successfully been applied to sparse vector prediction problems. We observe that it belongs to a wider class of norms, which we call the box-norms. Within this framework we derive an efficient alg…

Matrix Completion