paper-with-me

홈 › Papers

An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence

2026-02-02 · Qizhen Zhang, Ankush Garg, Jakob Foerster, Niladri Chatterji, Kshitiz Malik, Mike Lewis arxiv

Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM pretrainers often speculate that such noise contributes to instabilities in large-scale LLM pretraining and, in the worst cases, loss divergence, this phenomenon remains poorly understood.In this work, we present a systematic empirical study of whether noisy data causes LLM pretraining divergences and how it does so. By injecting controlled synthetic uniformly random noise into otherwise clean datasets, we analyze training dynamics across model sizes ranging from 480M to 5.2B parameters. We show that noisy data indeed induces training loss divergence, and that the probability of divergence depends strongly on the noise type, amount of noise, and model scale. We further find that noise-induced divergences exhibit activation patterns distinct from those caused by high learning rates, and we provide diagnostics that differentiate these two failure modes. Together, these results provide a large-scale, controlled characterization of how noisy data affects loss divergence in LLM pretraining.

📄 PDF Abstract BibTeX arXiv:2602.02400

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Distribution-Free Pretraining of Classification Losses via Evolutionary Dynamics

2026-05-05 · Meng Xiang, Yan Pei arxiv

We propose Evolutionary Dynamic Loss (EDL), a framework that learns a transferable classification loss in the probability space using unlimited synthetic prediction-label pairs, without accessing real samples during the …

Data Weighted Training Strategies for Grammatical Error Correction

2020-08-07 · Jared Lichtarge, Chris Alberti, Shankar Kumar

Recent progress in the task of Grammatical Error Correction (GEC) has been driven by addressing data sparsity, both through new methods for generating large and noisy pretraining data and through the publication of small…

Grammatical Error CorrectionMachine TranslationNMTTranslation

Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining

2025-09-30 · Anirudh Subramanyam, Yuxin Chen, Robert L. Grossman arxiv

Scaling laws for language model training traditionally characterize how performance scales with model size and dataset volume. Prior work has explored architecture variants and data treatments such as dataset filtering a…

Machine Translation

Learning with Noisy Labels

2013-12-01 · NeurIPS 2013 12 · Nagarajan Natarajan, Inderjit S. Dhillon, Pradeep K. Ravikumar, Ambuj Tewari

In this paper, we theoretically study the problem of binary classification in the presence of random classification noise --- the learner, instead of seeing the true labels, sees labels that have independently been flipp…

Binary ClassificationGeneral ClassificationLearning with noisy labels

Contrastive Visual-Linguistic Pretraining

2020-07-26 · Lei Shi, Kai Shuang, Shijie Geng, Peng Su 외

Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently. Such approaches can achieve superior performance due to the high-level semantic information captured durin…

Contrastive LearningregressionRepresentation LearningVisual Question Answering (VQA)