paper-with-me

Papers

Concavifiability and convergence: necessary and sufficient conditions for gradient descent analysis

2019-05-28 · Thulasi Tholeti, Sheetal Kalyani

Convergence of the gradient descent algorithm has been attracting renewed interest due to its utility in deep learning applications. Even as multiple variants of gradient descent were proposed, the assumption that the gradient of the objective is Lipschitz continuous remained an integral part of the analysis until recently. In this work, we look at convergence analysis by focusing on a property that we term as concavifiability, instead of Lipschitz continuity of gradients. We show that concavifiability is a necessary and sufficient condition to satisfy the upper quadratic approximation which is key in proving that the objective function decreases after every gradient descent update. We also show that any gradient Lipschitz function satisfies concavifiability. A constant known as the concavifier analogous to the gradient Lipschitz constant is derived which is indicative of the optimal step size. As an application, we demonstrate the utility of finding the concavifier the in convergence of gradient descent through an example inspired by neural networks. We derive bounds on the concavifier to obtain a fixed step size for a single hidden layer ReLU network.

📄 PDF Abstract BibTeX arXiv:1905.11620

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation

2025-05-06 · Max Qiushi Lin, Jincheng Mei, Matin Aghaei, Michael Lu 외

Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typically used with function approximation. In…

Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations

2024-03-12 · Akshay Kumar, Jarvis Haupt

This paper studies the gradient flow dynamics that arise when training deep homogeneous neural networks assumed to have locally Lipschitz gradients and an order of homogeneity strictly greater than two. It is shown here …

Targeted Separation and Convergence with Kernel Discrepancies

2022-09-26 · Alessandro Barp, Carl-Johann Simon-Gabriel, Mark Girolami, Lester Mackey

Maximum mean discrepancies (MMDs) like the kernel Stein discrepancy (KSD) have grown central to a wide range of applications, including hypothesis testing, sampler selection, distribution approximation, and variational i…

Variational Inference

Convergence of Deep ReLU Networks

2021-07-27 · Yuesheng Xu, Haizhang Zhang

We explore convergence of deep neural networks with the popular ReLU activation function, as the depth of the networks tends to infinity. To this end, we introduce the notion of activation domains and activation matrices…

image-classificationImage Classification

Analysis of Inter-Event Times in Linear Systems under Region-Based Self-Triggered Control

2022-12-29 · Anusree Rajan, Pavankumar Tallapragada

This paper analyzes the evolution of inter-event times (IETs) in linear systems under region-based self-triggered control (RBSTC). In this control method, the state space is partitioned into a finite number of conic regi…