paper-with-me

홈 › Papers

Towards Theoretical Understandings of Self-Consuming Generative Models

2024-02-19 · Shi Fu, Sen Zhang, Yingjie Wang, Xinmei Tian, DaCheng Tao

This paper tackles the emerging challenge of training generative models within a self-consuming loop, wherein successive generations of models are recursively trained on mixtures of real and synthetic data from previous generations. We construct a theoretical framework to rigorously evaluate how this training procedure impacts the data distributions learned by future models, including parametric and non-parametric models. Specifically, we derive bounds on the total variation (TV) distance between the synthetic data distributions produced by future models and the original real data distribution under various mixed training scenarios for diffusion models with a one-hidden-layer neural network score function. Our analysis demonstrates that this distance can be effectively controlled under the condition that mixed training dataset sizes or proportions of real data are large enough. Interestingly, we further unveil a phase transition induced by expanding synthetic data amounts, proving theoretically that while the TV distance exhibits an initial ascent, it declines beyond a threshold point. Finally, we present results for kernel density estimation, delivering nuanced insights such as the impact of mixed data training on error propagation.

📄 PDF Abstract BibTeX arXiv:2402.11778

Code (0)

등록된 구현이 없습니다.

Tasks

Density Estimation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Early Stopping Early Stopping is a regularization technique for deep neural networks that stops training when parameter updates no longer begin to yield improves on a validation set. In…

Similar Papers 제목 키워드 기반

Self-Correcting Self-Consuming Loops for Generative Model Training

2024-02-11 · Nate Gillman, Michael Freeman, Daksh Aggarwal, Chia-Hong Hsu 외

As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic…

Motion SynthesisRepresentation Learning

Self-Consuming Generative Models with Adversarially Curated Data

2025-05-14 · Xiukun Wei, Xueru Zhang

Recent advances in generative models have made it increasingly difficult to distinguish real data from model-generated synthetic data. Using synthetic data for successive training of future model generations creates "sel…

Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences

2024-06-12 · Damien Ferbach, Quentin Bertrand, Avishek Joey Bose, Gauthier Gidel

The rapid progress in generative models has resulted in impressive leaps in generation quality, blurring the lines between synthetic and real data. Web-scale datasets are now prone to the inevitable contamination by synt…

Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering

2025-11-16 · Zhongteng Cai, Yaxuan Wang, Yang Liu, Xueru Zhang arxiv

As synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a ``self-consuming loop" that can lead to training instability or \textit{model co…

Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation

2025-11-12 · Hongru Zhao, Jinwen Fu, Tuan Pham arxiv

Self-consuming generative models have received significant attention over the last few years. In this paper, we study a self-consuming generative model with heterogeneous preferences that is a generalization of the model…