paper-with-me

홈 › Papers

Approximate Muon with low-rank adapters

2026-08-14 · Ben Anson, Conor Houghton, Edward Milsom arxiv

The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since it is not mathematically possible to orthogonalize the weight update given by a low-rank parameterization. In this paper, we address this issue by approximating the solution to a relaxed Muon objective in the low-rank setting via linearization and then least-squares. We provide an efficient implementation that uses matmul operations only, as opposed to more complex linear algebra decomposition routines. Our method, sMuon (small Muon), performs favourably across SFT and a ReLoRA pretraining experiment. While results are model- and eval-dependent, we find overall that using Muon for low-rank fine-tuning provides moderate performance improvements.

📄 PDF Abstract BibTeX arXiv:2608.14492

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters

2025-07-16 · Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Denis Bobkov 외 arxiv

This work presents a novel, fully Riemannian framework for Low-Rank Adaptation (LoRA) that geometrically treats low-rank adapters by optimizing them directly on the fixed-rank manifold. This formulation eliminates the pa…

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training

2025-09-15 · Chuan He, Zhanwang Deng, Zhaosong Lu arxiv

Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlooked. Recently, the optimizer Muon \citep{jordanmuon}, which explicitly …

On the Convergence Analysis of Muon

2025-05-29 · Wei Shen, Ruichuan Huang, Minhui Huang, Cong Shen 외

The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors during optimization, potentially overlooking …

Reassessing Muon for Matrix Factorization

2026-07-14 · Ali Parviz, Gal Mishne, Alex Cloninger arxiv

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language…

LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold

2026-06-11 · Franz Louis Cesista, Katherine Crowson, Cédric Simal, Stella Biderman arxiv

Low-Rank Adaptation (LoRA) significantly reduces compute and memory costs for finetuning Deep Learning models but is often harder to tune than dense training: when using factor-wise optimizers such as AdamW, it is sensit…