paper-with-me

홈 › Papers

Instability in Downstream Task Performance During LLM Pretraining

2025-10-06 · Yuto Nishida, Masaru Isonuma, Yusuke Oda arxiv

When training large language models (LLMs), it is common practice to track downstream task performance throughout the training process and select the checkpoint with the highest validation score. However, downstream metrics often exhibit substantial fluctuations, making it difficult to identify the checkpoint that truly represents the best-performing model. In this study, we empirically analyze the stability of downstream task performance in an LLM trained on diverse web-scale corpora. We find that task scores frequently fluctuate throughout training, both at the aggregate and example levels. To address this instability, we investigate two post-hoc checkpoint integration methods: checkpoint averaging and ensemble, motivated by the hypothesis that aggregating neighboring checkpoints can reduce performance volatility. We demonstrate both empirically and theoretically that these methods improve downstream performance stability without requiring any changes to the training procedure.

📄 PDF Abstract BibTeX arXiv:2510.04848

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The SSL Interplay: Augmentations, Inductive Bias, and Generalization

2023-02-06 · Vivien Cabannes, Bobak T. Kiani, Randall Balestriero, Yann Lecun 외

Self-supervised learning (SSL) has emerged as a powerful framework to learn representations from raw data without supervision. Yet in practice, engineers face issues such as instability in tuning optimizers and collapse …

Data AugmentationInductive BiasSelf-Supervised Learning

Meta-learning for downstream aware and agnostic pretraining

2021-06-06 · Hongyin Luo, Shuyan Dong, Yung-Sung Chuang, Shang-Wen Li

Neural network pretraining is gaining attention due to its outstanding performance in natural language processing applications. However, pretraining usually leverages predefined task sequences to learn general linguistic…

Meta-Learning

Revisiting Theory of Contrastive Learning for Domain Generalization

2025-12-02 · Ali Alvandi, Mina Rezaei arxiv

Contrastive learning is among the most popular and powerful approaches for self-supervised representation learning, where the goal is to map semantically similar samples close together while separating dissimilar ones in…

Representation LearningDomain GeneralizationContrastive Learning

Difference-Masking: Choosing What to Mask in Continued Pretraining

2023-05-23 · Alex Wilf, Syeda Nahida Akter, Leena Mathur, Paul Pu Liang 외

The self-supervised objective of masking-and-predicting has led to promising performance gains on a variety of downstream tasks. However, while most approaches randomly mask tokens, there is strong intuition that decidin…

Self-Supervised Learning

Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima

2026-04-10 · Huanran Chen, Huaqing Zhang, Xiao Li, Yinpeng Dong 외 arxiv

The foundational capabilities of large language models are acquired during pretraining on internet-scale, highly heterogeneous data mixtures. In this work, we investigate an interesting geometric question regarding the c…