paper-with-me

홈 › Papers

Lotus: Efficient LLM Training by Randomized Low-Rank Gradient Projection with Adaptive Subspace Switching

2026-02-01 · Tianhao Miao, Zhongyuan Bao, Lejun Zhang arxiv

Training efficiency in large-scale models is typically assessed through memory consumption, training time, and model performance. Current methods often exhibit trade-offs among these metrics, as optimizing one generally degrades at least one of the others. Addressing this trade-off remains a central challenge in algorithm design. While GaLore enables memory-efficient training by updating gradients in a low-rank subspace, it incurs a comparable extra training time cost due to the Singular Value Decomposition(SVD) process on gradients. In this paper, we propose Lotus, a method that resolves this trade-off by simply modifying the projection process. We propose a criterion that quantifies the displacement of the unit gradient to enable efficient transitions between low-rank gradient subspaces. Experimental results indicate that Lotus is the most efficient method, achieving a 30% reduction in training time and a 40% decrease in memory consumption for gradient and optimizer states. Additionally, it outperforms the baseline method in both pre-training and fine-tuning tasks.

📄 PDF Abstract BibTeX arXiv:2602.01233

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Geometrically Principled Randomized Optimization for Efficient LLM Training

2025-10-02 · Sahar Rajabi, Nayeema Nonta, Sirisha Rambhatla arxiv

Low-rank gradient optimization for large language models is currently divided into two categories: structured methods that rigorously identify subspaces, and randomized approaches employed primarily for computational eff…

Computational Efficiency

AdaRankGrad: Adaptive Gradient-Rank and Moments for Memory-Efficient LLMs Training and Fine-Tuning

2024-10-23 · Yehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel 외

Training and fine-tuning large language models (LLMs) come with challenges related to memory and computational requirements due to the increasing size of the model weights and the optimizer states. Various techniques hav…

Structural Conditions for Projection-Cost Preservation via Randomized Matrix Multiplication

2017-05-29 · Agniva Chowdhury, Jiasen Yang, Petros Drineas

Projection-cost preservation is a low-rank approximation guarantee which ensures that the cost of any rank-$k$ projection can be preserved using a smaller sketch of the original data matrix. We present a general structur…

LORENZA: Enhancing Generalization in Low-Rank Gradient LLM Training via Efficient Zeroth-Order Adaptive SAM

2025-02-26 · Yehonathan Refael, Iftach Arbel, Ofir Lindenbaum, Tom Tirer

We study robust parameter-efficient fine-tuning (PEFT) techniques designed to improve accuracy and generalization while operating within strict computational and memory hardware constraints, specifically focusing on larg…

parameter-efficient fine-tuning

Low-Rank Matrix Estimation From Rank-One Projections by Unlifted Convex Optimization

2020-04-06 · Sohail Bahmani, Kiryung Lee

We study an estimator with a convex formulation for recovery of low-rank matrices from rank-one projections. Using initial estimates of the factors of the target $d_1\times d_2$ matrix of rank-$r$, the estimator admits a…