paper-with-me

홈 › Papers

Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods

2026-06-17 · Depen Morwani, Alexandru Meterez, Pranav Nair, Sham Kakade arxiv

Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al., 2018] are widely used in modern training, but their stochastic benefits depend on two distinct quantities: serial runtime, the number of iterations needed to reach a target accuracy, and compute efficiency (CE), the inverse total gradient-query or FLOP cost. Larger batches reduce serial runtime without hurting CE only when the contraction gap grows linearly with batch size. We study stochastic HB and ASGD for consistent linear regression with Gaussian covariates and prove finite-dimensional, discrete-time lower bounds on their batch-size tradeoffs. Our first result shows that HB does not improve the CE frontier over SGD for arbitrary spectra; rather, it preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale. This window can be a factor $\sqrtκ$ larger than the SGD critical batch size. For ASGD, the picture is more spectrum-dependent: for rapidly decaying power-law spectra, ASGD improves small-batch CE over HB/SGD, but as batch size grows it trades this CE advantage for improved serial runtime. Synthetic linear-regression experiments verify these qualitative regimes, including near-overlap of ASGD and HB for slowly decaying spectra and the predicted CE--serial tradeoff for rapidly decaying spectra.

📄 PDF Abstract BibTeX arXiv:2606.19179

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer Models

2023-09-21 · NeurIPS 2023 11

Large language models (LLMs) are highly capable but also computationally expensive. Characterizing the _fundamental tradeoff_ between inference efficiency and model capabilities requires a metric that is comparable acro…

Accurate Reaction-Diffusion Operator Splitting on Tetrahedral Meshes for Parallel Stochastic Molecular Simulations

2015-12-10

Spatial stochastic molecular simulations in biology are limited by the intense computation required to track molecules in space either in a discrete time or discrete space framework, meaning that the serial limit has alr…

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

2026-05-26 · Yipeng Ouyang, Xin Huang, Bingjie Liu, Zhongchun Zheng 외 arxiv

LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain largely centered on static, isolated, and short-horizon benchmarks th…

Learning to Inference Adaptively for Multimodal Large Language Models

2025-03-13 · Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi 외

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in reasoning, yet come with substantial computational cost, limiting their deployment in resource-constrained settings. Despite recent efforts o…

HallucinationQuestion Answering

Robust, fast and accurate: a 3-step method for automatic histological image registration

2019-03-28 · Johannes Lotz, Nick Weiss, Stefan Heldmann

We present a 3-step registration pipeline for differently stained histological serial sections that consists of 1) a robust pre-alignment, 2) a parametric registration computed on coarse resolution images, and 3) an accu…

Image Registration