paper-with-me

홈 › Papers

Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling

2026-05-29 · Dmitrii Feoktistov, Timofey Belinsky, Andrey Veprikov, Amir Zainullin, Aleksandr Beznosikov arxiv

Sign-based and LMO-inspired optimizers have recently attracted substantial attention in deep learning due to their strong performance and low memory footprint. However, their fixed-magnitude updates can hurt terminal convergence: they decouple update mechanisms from gradient magnitudes and fail to account for parameter heterogeneity, often leading to oscillation rather than convergence. We propose SoftSignum, a smooth relaxation of sign-based optimization that replaces the hard sign map with a temperature-controlled soft-sign transformation, enabling a parameter-wise transition from sign-like updates to magnitude-sensitive SGD-like steps. We complement it with an adaptive quantile-based temperature schedule and extend the same principle to matrix-valued optimizers, obtaining SoftMuon. We also develop a generalized geometry-relaxation framework based on strongly convex regularizers and Fenchel conjugates, proving convergence in stochastic non-convex setting. Experiments on diverse deep learning tasks, including LLM pretraining, show that SoftSignum and SoftMuon consistently improve over their hard sign-based counterparts and standard AdamW.

📄 PDF Abstract BibTeX arXiv:2605.31371

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Holonorm

2025-11-13 · Daryl Noupa Yongueng, Hamidou Tembine arxiv

Normalization is a key point in transformer training . In Dynamic Tanh (DyT), the author demonstrated that Tanh can be used as an alternative layer normalization (LN) and confirmed the effectiveness of the idea. But Tanh…

Hybrid activation functions for deep neural networks: S3 and S4 -- a novel approach to gradient flow optimization

2025-07-29 · Sergii Kavun arxiv

Activation functions are critical components in deep neural networks, directly influencing gradient flow, training stability, and model performance. Traditional functions like ReLU suffer from dead neuron problems, while…

Multi-class ClassificationBinary Classification

Re-parameterizing Your Optimizers rather than Architectures

2022-05-30 · Xiaohan Ding, Honghao Chen, Xiangyu Zhang, Kaiqi Huang 외

The well-designed structures in neural networks reflect the prior knowledge incorporated into the models. However, though different models have various priors, we are used to training them with model-agnostic optimizers …

Quantization

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam

2025-07-09 · Hanyang Peng, Shuang Qin, Yue Yu, Fangqing Jiang 외 arxiv

Adam has proven remarkable successful in training deep neural networks, but the mechanisms underlying its empirical successes and limitations remain underexplored. In this study, we demonstrate that the effectiveness of …

Stochastic Optimization

Is Hyper-Parameter Optimization Different for Software Analytics?

2024-01-17 · Rahul Yedida, Tim Menzies

Yes. SE data can have "smoother" boundaries between classes (compared to traditional AI data sets). To be more precise, the magnitude of the second derivative of the loss function found in SE data is typically much small…