paper-with-me

Papers

An Exponential Averaging Process with Strong Convergence Properties

2025-05-15 · Frederik Köhne, Anton Schiela

Averaging, or smoothing, is a fundamental approach to obtain stable, de-noised estimates from noisy observations. In certain scenarios, observations made along trajectories of random dynamical systems are of particular interest. One popular smoothing technique for such a scenario is exponential moving averaging (EMA), which assigns observations a weight that decreases exponentially in their age, thus giving younger observations a larger weight. However, EMA fails to enjoy strong stochastic convergence properties, which stems from the fact that the weight assigned to the youngest observation is constant over time, preventing the noise in the averaged quantity from decreasing to zero. In this work, we consider an adaptation to EMA, which we call $p$-EMA, where the weights assigned to the last observations decrease to zero at a subharmonic rate. We provide stochastic convergence guarantees for this kind of averaging under mild assumptions on the autocorrelations of the underlying random dynamical system. We further discuss the implications of our results for a recently introduced adaptive step size control for Stochastic Gradient Descent (SGD), which uses $p$-EMA for averaging noisy observations.

📄 PDF Abstract BibTeX arXiv:2505.10605

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exponential integrability properties of Euler discretization schemes for the Cox-Ingersoll-Ross process

2015-12-17

We analyze exponential integrability properties of the Cox-Ingersoll-Ross (CIR) process and its Euler discretizations with various types of truncation and reflection at 0. These properties play a key role in establishing…

Second-Order Mirror Descent: Convergence in Games Beyond Averaging and Discounting

2021-11-18 · Bolin Gao, Lacra Pavel

In this paper, we propose a second-order extension of the continuous-time game-theoretic mirror descent (MD) dynamics, referred to as MD2, which provably converges to mere (but not necessarily strict) variationally stabl…

Averaging on the Bures-Wasserstein manifold: dimension-free convergence of gradient descent

2021-06-16 · NeurIPS 2021 12 · Jason M. Altschuler, Sinho Chewi, Patrik Gerber, Austin J. Stromme

We study first-order optimization algorithms for computing the barycenter of Gaussian distributions with respect to the optimal transport metric. Although the objective is geodesically non-convex, Riemannian GD empirical…

Gradient Descent Averaging and Primal-dual Averaging for Strongly Convex Optimization

2020-12-29 · Wei Tao, Wei Li, Zhisong Pan, Qing Tao

Averaging scheme has attracted extensive attention in deep learning as well as traditional machine learning. It achieves theoretically optimal convergence and also improves the empirical model performance. However, there…

Stochastic Gradient Descent with Exponential Convergence Rates of Expected Classification Errors

2018-06-14 · Atsushi Nitanda, Taiji Suzuki

We consider stochastic gradient descent and its averaging variant for binary classification problems in a reproducing kernel Hilbert space. In the traditional analysis using a consistency property of loss functions, it i…

Binary ClassificationClassificationGeneral Classification