paper-with-me

홈 › Papers

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

2026-08-06 · Iosif Lytras, Nikolaos Makras, Sotirios Sabanis arxiv

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.

📄 PDF Abstract BibTeX arXiv:2608.06283

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Error estimates for tamed Euler and Randomized Euler schemes for SDEs with locally Lipschitz drift with applications to non-logconcave sampling and optimization

2026-05-24 · Iosif Lytras, Angelos Ntousis arxiv

In this paper, we study the numerical discretization of stochastic differential equations with locally Lipschitz, super-linearly growing drift, and the resulting implications for sampling from non-log-concave distributio…

Taming neural networks with TUSLA: Non-convex learning via adaptive stochastic gradient Langevin algorithms

2020-06-25 · Attila Lovas, Iosif Lytras, Miklós Rásonyi, Sotirios Sabanis

Artificial neural networks (ANNs) are typically highly nonlinear systems which are finely tuned via the optimization of their associated, non-convex loss functions. In many cases, the gradient of any such loss function h…

Tamed Stochastic Gradient Hamiltonian Monte Carlo

2026-07-16 · Zhuoran Wang, Ying Zhang arxiv

In this paper, we propose a novel tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) algorithm for sampling and stochastic optimization problems with superlinearly growing stochastic gradients. Under a certain co…

Stochastic Optimization

Polygonal Unadjusted Langevin Algorithms: Creating stable and efficient adaptive algorithms for neural networks

2021-05-28 · Dong-Young Lim, Sotirios Sabanis

We present a new class of Langevin based algorithms, which overcomes many of the known shortcomings of popular adaptive optimizers that are currently used for the fine tuning of deep learning models. Its underpinning the…

Deep LearningStochastic Optimization

Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin

2026-07-16 · Yifan Chen, Xiaoou Cheng, Jonathan Niles-Weed, Jonathan Weare arxiv

Unadjusted samplers such as unadjusted Hamiltonian Monte Carlo and underdamped Langevin are well-known to be biased. Metropolis--Hastings adjustment has been conventionally incorporated into Hamiltonian Monte Carlo to el…