paper-with-me

홈 › Papers

LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters

2025-07-16 · Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Denis Bobkov, Vera Soboleva, Aibek Alanov, Maxim Rakhuba arxiv

This work presents a novel, fully Riemannian framework for Low-Rank Adaptation (LoRA) that geometrically treats low-rank adapters by optimizing them directly on the fixed-rank manifold. This formulation eliminates the parametrization ambiguity present in standard Euclidean optimizers. Our framework integrates three key components to achieve this: (1) we derive Riemannion, a new Riemannian optimizer on the fixed-rank matrix manifold that generalizes the recently proposed Muon optimizer; (2) we develop a Riemannian gradient-informed LoRA initialization, and (3) we provide an efficient implementation without prominent overhead that uses automatic differentiation to compute arising geometric operations while adhering to best practices in numerical linear algebra. Comprehensive experimental results on both LLM and diffusion model architectures demonstrate that our approach yields consistent and noticeable improvements in convergence speed and final task performance over both standard LoRA and its state-of-the-art modifications.

📄 PDF Abstract BibTeX arXiv:2507.12142

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold

2026-06-11 · Franz Louis Cesista, Katherine Crowson, Cédric Simal, Stella Biderman arxiv

Low-Rank Adaptation (LoRA) significantly reduces compute and memory costs for finetuning Deep Learning models but is often harder to tune than dense training: when using factor-wise optimizers such as AdamW, it is sensit…

MARS-M: When Variance Reduction Meets Matrices

2025-10-20 · Yifeng Liu, Angela Yuan, Quanquan Gu arxiv

Matrix-based preconditioned optimizers, such as Muon, have recently been shown to be more efficient than scalar-based optimizers for training large-scale neural networks, including large language models (LLMs). Recent be…

Can Muon Fine-tune Adam-Pretrained Models?

2026-05-11 · Xingyu Qu, Peigeng Huang, Samuel Horvath arxiv

Muon has emerged as an efficient alternative to Adam for pretraining, yet remains underused for fine-tuning. A key obstacle is that most open models are pretrained with Adam, and naively switching to Muon for fine-tuning…

When Muon Optimizer Meets Adversarial Training: A Theoretical and Empirical Study

2026-05-26 · Jun Yan, Weiquan Huang, Jiankai Zuo, Yujian Mo 외 arxiv

Adversarial training (AT) remains one of the most reliable empirical defenses against adversarial attacks. Its robustness critically depends on how the underlying min-max objective is optimized. In practice, Stochastic G…

Approximate Muon with low-rank adapters

2026-08-14 · Ben Anson, Conor Houghton, Edward Milsom arxiv

The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common P…

parameter-efficient fine-tuning