Reformulation for Pretraining Data Augmentation
Despite the impressive capabilities of large language models across various tasks, their continued scaling is severely hampered not only by data scarcity but also by the performance degradation associated with excessive data repetition during training. To overcome this critical bottleneck, we propose the Massive Genre-Audience(MGA) reformulation method, a lightweight and scalable data augmentation technique inspired by synthetic data methodologies. MGA systematically reformulates existing corpora into diverse, contextually-rich variations to mitigate the negative effects of repetition, and we introduce this approach along with the resulting 770 billion token MGACorpus in this work. We experimentally validate its core benefit by demonstrating superior performance against data repetition and upsampling in scaling scenarios (up to 13B parameters). Furthermore, comprehensive analysis investigates the role of prompt engineering in generation quality and reveals nuances in evaluating model capabilities using standard loss metrics. Our work shows that MGA provides a reliable pathway to substantially augment training datasets, effectively alleviating repetition bottlenecks and enabling more efficient scaling of large language models.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationPrompt EngineeringSimilar Papers 제목 키워드 기반
Answer-Supervised Question Reformulation for Enhancing Conversational Machine Comprehension
In conversational machine comprehension, it has become one of the research hotspots integrating conversational history information through question reformulation for obtaining better answers. However, the existing questi…
Reading Comprehensionreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1PAITS: Pretraining and Augmentation for Irregularly-Sampled Time Series
Real-world time series data that commonly reflect sequential human behavior are often uniquely irregularly sampled and sparse, with highly nonuniform sampling over time and entities. Yet, commonly-used pretraining and au…
Time SeriesDemystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands prod…
Data AugmentationText GenerationSimple and Effective Input Reformulations for Translation
Foundation language models learn from their finetuning input context in different ways. In this paper, we reformulate inputs during finetuning for challenging translation tasks, leveraging model strengths from pretrainin…
TranslationConnect Later: Improving Fine-tuning for Robustness with Targeted Augmentations
Models trained on a labeled source domain (e.g., labeled images from wildlife camera traps) often generalize poorly when deployed on an out-of-distribution (OOD) target domain (e.g., images from new camera trap locations…
Contrastive LearningDomain AdaptationTime SeriesTime Series Classification