paper-with-me

홈 › Papers

Scaling and Transferability of Annealing Strategies in Large Language Model Training

2025-12-05 · Siqi Wang, Zhengyu Chen, Teng Xiao, Zheqi Lv, Jinluan Yang, Xunliang Cai, Jingang Wang, Xiaomeng Li arxiv

Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a generalized predictive framework for optimizing annealing strategies under the Warmup-Steady-Decay (WSD) scheduler. Our improved framework incorporates training steps, maximum learning rate, and annealing behavior, enabling more efficient optimization of learning rate schedules. Our work provides a practical guidance for selecting optimal annealing strategies without exhaustive hyperparameter searches, demonstrating that smaller models can serve as reliable proxies for optimizing the training dynamics of larger models. We validate our findings on extensive experiments using both Dense and Mixture-of-Experts (MoE) models, demonstrating that optimal annealing ratios follow consistent patterns and can be transferred across different training configurations.

📄 PDF Abstract BibTeX arXiv:2512.13705

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling Law with Learning Rate Annealing

2024-08-20 · Howe Tissue, Venus Wang, Lu Wang

We find that the cross-entropy loss curves of neural language models empirically adhere to a scaling law with learning rate (LR) annealing over training steps: $$L(s) = L_0 + A\cdot S_1^{-\alpha} - C\cdot S_2,$$ where $L…

Language Modelling

Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models

2024-10-08 · Siqi Wang, Zhengyu Chen, Bei Li, Keqing He 외

The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferability and discrepancies of scaling laws b…

Mixture-of-Experts

Scaling Laws for Black box Adversarial Attacks

2024-11-25 · Chuan Liu, Huanran Chen, Yichi Zhang, Yinpeng Dong 외

Adversarial examples usually exhibit good cross-model transferability, enabling attacks on black-box models with limited information about their architectures and parameters, which are highly threatening in commercial bl…

Adversarial Attack

An Investigation on Hardware-Aware Vision Transformer Scaling

2021-09-29 · Chaojian Li, KyungMin Kim, Bichen Wu, Peizhao Zhang 외

Vision Transformer (ViT) has demonstrated promising performance in various computer vision tasks, and recently attracted a lot of research attention. Many recent works have focused on proposing new architectures to impro…

GPUimage-classificationImage Classificationobject-detection+2

Towards Transferable Speech Emotion Representation: On loss functions for cross-lingual latent representations

2022-03-28 · Sneha Das, Nicole Nadine Lønfeldt, Anne Katrine Pagsberg, Line H. Clemmensen

In recent years, speech emotion recognition (SER) has been used in wide ranging applications, from healthcare to the commercial sector. In addition to signal processing approaches, methods for SER now also use deep learn…

ClassificationDenoisingEmotion ClassificationEmotion Recognition+2