paper-with-me

Papers

Entropic Regularization in the Deep Linear Network

2025-12-05 · Alan Chen, Tejas Kotwal, Govind Menon arxiv

We study regularization for the deep linear network (DLN) using the entropy formula introduced in arXiv:2509.09088. The equilibria and gradient flow of the free energy on the Riemannian manifold of end-to-end maps of the DLN are characterized for energies that depend symmetrically on the singular values of the end-to-end matrix. The only equilibria are minimizers and the set of minimizers is an orbit of the orthogonal group. In contrast with random matrix theory there is no singular value repulsion. The corresponding gradient flow reduces to a one-dimensional ordinary differential equation whose solution gives explicit relaxation rates toward the minimizers. We also study the concavity of the entropy in the chamber of singular values. The entropy is shown to be strictly concave in the Euclidean geometry on the chamber but not in the Riemannian geometry defined by the DLN metric.

📄 PDF Abstract BibTeX arXiv:2512.06137

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toric Geometry of Entropic Regularization

2022-02-03 · Bernd Sturmfels, Simon Telen, François-Xavier Vialard, Max von Renesse

Entropic regularization is a method for large-scale linear programming. Geometrically, one traces intersections of the feasible polytope with scaled toric varieties, starting at the Birch point. We compare this to log-ba…

A Sinkhorn-type Algorithm for Constrained Optimal Transport

2024-03-08 · Xun Tang, Holakou Rahmanian, Michael Shavlovsky, Kiran Koshy Thekumparampil 외

Entropic optimal transport (OT) and the Sinkhorn algorithm have made it practical for machine learning practitioners to perform the fundamental task of calculating transport distance between statistical distributions. In…

Scheduling

Understanding Entropic Regularization in GANs

2021-11-02 · Daria Reshetova, Yikun Bai, Xiugang Wu, Ayfer Ozgur

Generative Adversarial Networks are a popular method for learning distributions from data by modeling the target distribution as a function of a known distribution. The function, often referred to as the generator, is op…

The rate of convergence of Bregman proximal methods: Local geometry vs. regularity vs. sharpness

2022-11-15 · Waïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis Mertikopoulos

We examine the last-iterate convergence rate of Bregman proximal methods - from mirror descent to mirror-prox and its optimistic variants - as a function of the local geometry induced by the prox-mapping defining the met…

Extended Formulations for Online Linear Bandit Optimization

2013-11-20 · Shaona Ghosh, Adam Prugel-Bennett

On-line linear optimization on combinatorial action sets (d-dimensional actions) with bandit feedback, is known to have complexity in the order of the dimension of the problem. The exponential weighted strategy achieves …

Efficient Exploration