paper-with-me

홈 › Papers

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

2026-08-21 · Guoxiang Xu, Bince Qu, Qi Sun, Cheng Zhuo arxiv

Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight, jointly changing its norm and direction. This interaction matters because the current norm determines angular motion, while directional learning can drive norm growth and thereby alter later steps. We introduce RODE, which gives the radial and directional components separate update rules and step sizes. RODE explicitly updates the matrix Frobenius norm through a scalar radial rule, while its directional channel performs Newton--Schulz-conditioned updates in the tangent space. Controlled GPT-2 interventions show gains from both direct norm control and RODE's directional update. Across two language-modeling and two image-classification tasks, RODE outperforms both Muon variants in every direct comparison and ends with lower full-model norms. At 1.5B scale, using the learning rate transferred directly from the Qwen2-style LM sweep, RODE lowers loss from 4.145 to 3.346 and final global norm from 11964 to 2183 relative to Muon RMS, with fixed-radius RODE improving further. For Qwen3.5-9B full-parameter fine-tuning, all six optimizers use the same tuning budget and the same formal-training and evaluation settings; RODE outperforms both Muon variants on all four evaluation tasks and attains the highest mean on GSM8K and MATH-500. Thus, decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.

📄 PDF Abstract BibTeX arXiv:2608.21024

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decoupled Orthogonal Dynamics: Regularization for Deep Network Optimizers

2026-02-04 · Hao Chen, Jinghui Yuan, Hanmin Zhang arxiv

Is the standard weight decay in AdamW truly optimal? Although AdamW decouples weight decay from adaptive gradient scaling, a fundamental conflict remains: the Radial Tug-of-War. In deep learning, gradients tend to increa…

Soft-Radial Projection for Constrained End-to-End Learning

2026-02-03 · Philipp J. Schneider, Daniel Kuhn arxiv

Integrating hard constraints into deep learning is essential for safety-critical systems. Yet existing constructive layers that project predictions onto constraint boundaries face a fundamental bottleneck: gradient satur…

Cooperative Target Capture in 3D Engagements over Switched Dynamic Graphs

2025-07-02 · Abhinav Sinha, Shashi Ranjan Kumar arxiv

This paper presents a leaderless cooperative guidance strategy for simultaneous time-constrained interception of a stationary target when the interceptors exchange information over switched dynamic graphs. We specificall…

Universal approximation and model compression for radial neural networks

2021-07-06 · Iordan Ganev, Twan van Laarhoven, Robin Walters

We introduce a class of fully-connected neural networks whose activation functions, rather than being pointwise, rescale feature vectors by a function depending only on their norm. We call such networks radial neural net…

Model Compression

AeDet: Azimuth-invariant Multi-view 3D Object Detection

2022-11-22 · CVPR 2023 1 · Chengjian Feng, Zequn Jie, Yujie Zhong, Xiangxiang Chu 외

Recent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry…

3D Object DetectionDepth EstimationDepth PredictionObject+2