paper-with-me

홈 › Papers

Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA

2026-02-10 · Sangyoon Lee, Jaeho Lee arxiv

Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical gains, often on the same benchmarks. We show that these contradictions arise from a single overlooked factor: the batch size. When properly tuned, vanilla LoRA often matches the performance of more complex variants. We further propose a proxy-based, cost-efficient strategy for batch size tuning, revealing the impact of rank, dataset size, and model capacity on the optimal batch size. Our findings elevate batch size from a minor implementation detail to a first-order design parameter, reconciling prior inconsistencies and enabling more reliable evaluations of LoRA variants.

📄 PDF Abstract BibTeX arXiv:2602.09492

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inefficiency of K-FAC for Large Batch Size Training

2019-03-14 · Linjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao 외

In stochastic optimization, using large batch sizes during training can leverage parallel resources to produce faster wall-clock training times per training epoch. However, for both training loss and testing error, recen…

Stochastic Optimization

The Effect of Mini-Batch Noise on the Implicit Bias of Adam

2026-02-02 · Matias D. Cattaneo, Boris Shigida arxiv

With limited high-quality data and growing compute, multi-epoch training is gaining back its importance across sub-areas of deep learning. Adam(W), versions of which are go-to optimizers for many tasks such as next token…

DeepRacer on Physical Track: Parameters Exploration and Performance Evaluation

2024-06-06 · Sinan Koparan, Bahman Javadi

This paper focuses on the physical racetrack capabilities of AWS DeepRacer. Two separate experiments were conducted. The first experiment (Experiment I) focused on evaluating the impact of hyperparameters on the physical…

Object

Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful

2025-07-09 · Martin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson 외 arxiv

Conventional wisdom dictates that small batch sizes make language model pretraining and fine-tuning unstable, motivating gradient accumulation, which trades off the number of optimizer steps for a proportional increase i…

Critical Bach Size Minimizes Stochastic First-Order Oracle Complexity of Deep Learning Optimizer using Hyperparameters Close to One

2022-08-21 · Hideaki Iiduka

Practical results have shown that deep learning optimizers using small constant learning rates, hyperparameters close to one, and large batch sizes can find the model parameters of deep neural networks that minimize the …