paper-with-me

홈 › Papers

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers

2026-05-18 · Tim Tsz-Kit Lau, Weijie Su arxiv

A striking geometric disparity has long persisted in the practice of deep learning. While modern neural network architectures naturally exhibit rich symmetry and equivariance properties, popular optimizers such as Adam and its variants operate inherently coordinate-wise, rendering them unable to respect the equivariance structures of the parameter space. We address this disparity by introducing a symmetry-compatible principle for optimizer design: the gradient update rule should be equivariant under the symmetry group acting on the corresponding weight block. Following this principle, we first provide a unified perspective on bi-orthogonally equivariant updates for general matrix layers, as employed by stochastic spectral descent, Muon, Scion, and polar gradient methods. More importantly, by moving from orthogonal groups to permutation and shared-shift symmetries, we derive symmetry-compatible optimizers for parameter blocks whose symmetries differ from those of general matrix layers: embedding and LM head matrices, SwiGLU MLP projections, and MoE router matrices. These constructions include one-sided spectral, row-norm, hybrid row-norm/spectral, row-aware, column-aware, centered row-norm, and left-spectral updates. They yield an end-to-end layerwise optimizer stack in which each major matrix-valued parameter class is assigned an update whose equivariance matches its symmetry group. We corroborate this principle through pre-training experiments on dense and sparse MoE language models, including Qwen3-0.6B-style, Gemma 3 1B-style, OLMoE-1B-7B-style, and downsized gpt-oss architectures. Across these experiments, symmetry-compatible update rules consistently improve final validation loss, reduce expert load imbalance in sparse MoE models, and in several cases control final vocabulary-logit growth, improve router stability, and overall training stability over the corresponding AdamW updates.

📄 PDF Abstract BibTeX arXiv:2605.18106

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gauge-Equivariant Graph Neural Networks for Lattice Gauge Theories

2026-04-22 · Ali Rayat, Yaohang Li, Gia-Wei Chern arxiv

Local gauge symmetry underlies fundamental interactions and strongly correlated quantum matter, yet existing machine-learning approaches lack a general, principled framework for learning under site-dependent symmetries, …

Graph Neural Network

ARO: A New Lens On Matrix Optimization For Large Models

2026-02-09 · Wenbo Gong, Javier Zazo, Qijun Luo, Puqian Wang 외 arxiv

Matrix-based optimizers have attracted growing interest for improving LLM training efficiency, with significant progress centered on orthogonalization/whitening based methods. While yielding substantial performance gains…

Uncertainty Principle based optimization; new metaheuristics framework

2020-06-02 · Mojtaba Moattari, Mohammad Hassan Moradi, Emad Roshandel

To more flexibly balance between exploration and exploitation, a new meta-heuristic method based on Uncertainty Principle concepts is proposed in this paper. UP is is proved effective in multiple branches of science. In …

Designing color symmetry in stigmergic art

2021-06-28 · Hendrik Richter

Color symmetry is an extension of symmetry imposed by isometric transformations and means that the colors of geometrical objects are assigned according to the symmetry properties of the objects. A color symmetry permutes…

Neuro-evolutionary stochastic architectures in gauge-covariant neural fields

2026-04-22 · Rodrigo Carmo Terin arxiv

We extend our gauge-covariant stochastic neural-field framework by promoting architecture-level parameters to slow stochastic variables evolving in function space. Our effective theory is formulated in terms of classical…