paper-with-me

Papers

Schattor: Schatten-family methods for deep learning optimization

2026-06-14 · Bohao Ma, Junyu Zhang, Chuan He arxiv

Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis. Motivated by the limitations of SGD and the success of adaptive optimizers, we propose {\it Schattor}, a family of adaptive first-order methods based on Schatten norms. Schattor unifies SGD and the recently proposed matrix-variate adaptive optimizer Muon within a single Schatten-norm-based framework. We establish dimension-free stationarity guarantees for methods in the Schattor family for stochastic matrix optimization problems via a novel matrix martingale moment bound. We also develop multi-block extensions that adaptively balance block-wise optimization progress and prove dimension-free stationarity guarantees in this more general setting.

📄 PDF Abstract BibTeX arXiv:2606.15702

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unified Scalable Equivalent Formulations for Schatten Quasi-Norms

2016-06-02 · Fanhua Shang, Yuanyuan Liu, James Cheng

The Schatten quasi-norm can be used to bridge the gap between the nuclear norm and rank function, and is the tighter approximation to matrix rank. However, most existing Schatten quasi-norm minimization (SQNM) algorithms…

Simplifying Optimal Transport through Schatten-$p$ Regularization

2025-10-13 · Tyler Maunu arxiv

We propose a new general framework for recovering low-rank structure in optimal transport using Schatten-$p$ norm regularization. Our approach extends existing methods that promote sparse and interpretable transport maps…

Muon is Not That Special: Random or Inverted Spectra Work Just as Well

2026-05-11 · Zakhar Shumaylov, Nathaël Da Costa, Peter Zaika, Bálint Mucsányi 외 arxiv

The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear minimization oracle (LMO) theory. In this p…

Convex Tensor Decomposition via Structured Schatten Norm Regularization

2013-03-26 · NeurIPS 2013 12 · Ryota Tomioka, Taiji Suzuki

We discuss structured Schatten norms for tensor decomposition that includes two recently proposed norms ("overlapped" and "latent") for convex-optimization-based tensor decomposition, and connect tensor decomposition wit…

Tensor Decomposition

Learning Schatten--von Neumann Operators

2019-01-29 · Puoya Tabaghi, Maarten de Hoop, Ivan Dokmanić

We study the learnability of a class of compact operators known as Schatten--von Neumann operators. These operators between infinite-dimensional function spaces play a central role in a variety of applications in learnin…

Learning Theory