paper-with-me

홈 › Papers

Representing smooth functions as compositions of near-identity functions with implications for deep network optimization

2018-04-13 · Peter L. Bartlett, Steven N. Evans, Philip M. Long

We show that any smooth bi-Lipschitz $h$ can be represented exactly as a composition $h_m \circ ... \circ h_1$ of functions $h_1,...,h_m$ that are close to the identity in the sense that each $\left(h_i-\mathrm{Id}\right)$ is Lipschitz, and the Lipschitz constant decreases inversely with the number $m$ of functions composed. This implies that $h$ can be represented to any accuracy by a deep residual network whose nonlinear layers compute functions with a small Lipschitz constant. Next, we consider nonlinear regression with a composition of near-identity nonlinear maps. We show that, regarding Fr\'echet derivatives with respect to the $h_1,...,h_m$, any critical point of a quadratic criterion in this near-identity region must be a global minimizer. In contrast, if we consider derivatives with respect to parameters of a fixed-size residual network with sigmoid activation functions, we show that there are near-identity critical points that are suboptimal, even in the realizable case. Informally, this means that functional gradient methods for residual networks cannot get stuck at suboptimal critical points corresponding to near-identity layers, whereas parametric gradient methods for sigmoidal residual networks suffer from suboptimal critical points in the near-identity region.

📄 PDF Abstract BibTeX arXiv:1804.05012

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음

Similar Papers 제목 키워드 기반

Riemannian Stochastic Approximation for Minimizing Tame Nonsmooth Objective Functions

2023-02-01 · Johannes Aspman, Vyacheslav Kungurtsev, Reza Roohi Seraji

In many learning applications, the parameters in a model are structurally constrained in a way that can be modeled as them lying on a Riemannian manifold. Riemannian optimization, wherein procedures to enforce an iterati…

Riemannian optimization

Robust Basis Spline Decoupling for the Compression of Transformer Models

2026-05-11 · Joppe De Jonghe, Van Tien Pham, Mariya Ishteva arxiv

Decoupling is a powerful modeling paradigm for representing multivariate functions as compositions of linear transformations and univariate nonlinear functions. A single-layer decoupling can be viewed as a fully connecte…

Neural Network CompressionModel Compression

Graphical Convergence of Subgradients in Nonconvex Optimization and Learning

2018-10-17 · Damek Davis, Dmitriy Drusvyatskiy

We investigate the stochastic optimization problem of minimizing population risk, where the loss defining the risk is assumed to be weakly convex. Compositions of Lipschitz convex functions with smooth maps are the prima…

regressionStochastic Optimization

Step and Smooth Decompositions as Topological Clustering

2023-11-09 · Luciano Vinas, Arash A. Amini

We investigate a class of recovery problems for which observations are a noisy combination of continuous and step functions. These problems can be seen as non-injective instances of non-linear ICA with direct application…

Clustering

Bias-variance decompositions: the exclusive privilege of Bregman divergences

2025-01-30 · Tom Heskes

Bias-variance decompositions are widely used to understand the generalization performance of machine learning models. While the squared error loss permits a straightforward decomposition, other loss functions - such as z…