paper-with-me

홈 › Papers

Will we run out of data? Limits of LLM scaling based on human-generated data

2022-10-26 · Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, Marius Hobbhahn

We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock of public human text data. Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained. We explore how progress in language modeling can continue when human-generated text datasets cannot be scaled any further. We argue that synthetic data generation, transfer learning from data-rich domains, and data efficiency improvements might support further progress.

📄 PDF Abstract BibTeX arXiv:2211.04325

Code (1)

epoch-research/data-stock 공식 구현

Tasks

Language ModelingLanguage ModellingSynthetic Data GenerationTransfer Learning

Similar Papers 제목 키워드 기반

Experience Scaling: Post-Deployment Evolution For Large Language Models

2025-09-23 · Xingkun Yin, Kaibin Huang, Dong In Kim, Hongyang Du arxiv

Scaling model size, training data, and compute power have driven advances in large language models (LLMs), but these approaches are reaching saturation as human-generated text is exhausted and further gains diminish. We …

The Planetary Cost of AI Acceleration, Part II: The 10th Planetary Boundary and the 6.5-Year Countdown

2026-04-03 · William Yicheng Zhu, Lei Zhu arxiv

The recent, super-exponential scaling of autonomous Large Language Model (LLM) agents signals a broader, fundamental paradigm shift from machines primarily replacing the human hands (manual labor and mechanical processin…

A Tale of Tails: Model Collapse as a Change of Scaling Laws

2024-02-10 · Elvis Dohmatob, Yunzhen Feng, Pu Yang, Francois Charton 외

As AI model size grows, neural scaling laws have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data. Yet, the widespread u…

Language ModelingLanguage ModellingLarge Language ModelText Generation

PaSS: Parallel Speculative Sampling

2023-11-22 · Giovanni Monea, Armand Joulin, Edouard Grave

Scaling the size of language models to tens of billions of parameters has led to impressive performance on a wide range of tasks. At generation, these models are used auto-regressively, requiring a forward pass for each …

Evidence of a log scaling law for political persuasion with large language models

2024-06-20 · Kobi Hackenburg, Ben M. Tappin, Paul Röttger, Scott Hale 외

Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persu…

Persuasiveness