paper-with-me

홈 › Papers

Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training

2026-02-17 · Kevin Wang, Hongqian Niu, Didong Li arxiv

As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive degradation or collapse. In this article, we provide the first positive, rigorous theoretical results, to the best of our knowledge, showing that under model-agnostic mild conditions, the model converges to the true data-generating distribution. The convergence rate is the minimum of the model's intrinsic rate and the fraction of real data at each training iteration, revealing a phase transition between data-limited and model-limited regimes. We further show that, for biased real data, correcting the bias prevents the persistence and amplification of early bias over training iteration. Extensive experiments across simulations, real images and texts validate our theoretical framework, establishing quantitative conditions for long-term AI stability in contaminated environments.

📄 PDF Abstract BibTeX arXiv:2602.16065

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Synergetic Learning Systems: Concept, Architecture, and Algorithms

2020-05-31 · Ping Guo, Qian Yin

Drawing on the idea that brain development is a Darwinian process of ``evolution + selection'' and the idea that the current state is a local equilibrium state of many bodies with self-organization and evolution processe…

Decision Making

Rethinking the effects of data contamination in Code Intelligence

2025-06-03 · Zhen Yang, Hongyi Lin, Yifan He, Jie Xu 외

In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread adoption of Pretrained Language Models (PLMs) and Large Language Models (LLMs)…

Code GenerationCode SummarizationCode Translation

Biologically Inspired Design Concept Generation Using Generative Pre-Trained Transformers

2022-12-26 · Qihao Zhu, Xinyu Zhang, Jianxi Luo

Biological systems in nature have evolved for millions of years to adapt and survive the environment. Many features they developed can be inspirational and beneficial for solving technical problems in modern industries. …

Language Modelling

Death and Suicide in Universal Artificial Intelligence

2016-06-02 · Jarryd Martin, Tom Everitt, Marcus Hutter

Reinforcement learning (RL) is a general paradigm for studying intelligent behaviour, with applications ranging from artificial intelligence to psychology and economics. AIXI is a universal solution to the RL problem; it…

Reinforcement LearningReinforcement Learning (RL)

Generative Pre-Trained Transformers for Biologically Inspired Design

2022-03-31 · Qihao Zhu, Xinyu Zhang, Jianxi Luo

Biological systems in nature have evolved for millions of years to adapt and survive the environment. Many features they developed can be inspirational and beneficial for solving technical problems in modern industries. …

Language Modelling