paper-with-me

Papers

Model Collapse in the Self-Consuming Chain of Diffusion Finetuning: A Novel Perspective from Quantitative Trait Modeling

2024-07-04 · Youngseok Yoon, Dainong Hu, Iain Weissburg, Yao Qin, Haewon Jeong

The success of generative models has reached a unique threshold where their outputs are indistinguishable from real data, leading to the inevitable contamination of future data collection pipelines with synthetic data. While their potential to generate infinite samples initially offers promise for reducing data collection costs and addressing challenges in data-scarce fields, the severe degradation in performance has been observed when iterative loops of training and generation occur -- known as `model collapse.'' This paper explores a practical scenario in which a pretrained text-to-image diffusion model is finetuned using synthetic images generated from a previous iteration, a process we refer to as the `Chain of Diffusion.'' We first demonstrate the significant degradation in image quality caused by this iterative process and identify the key factor driving this decline through rigorous empirical investigations. Drawing an analogy between the Chain of Diffusion and biological evolution, we then introduce a novel theoretical analysis based on quantitative trait modeling. Our theoretical analysis aligns with empirical observations of the generated images in the Chain of Diffusion. Finally, we propose Reusable Diffusion Finetuning (ReDiFine), a simple yet effective strategy inspired by genetic mutations. ReDiFine mitigates model collapse without requiring any hyperparameter tuning, making it a plug-and-play solution for reusable image generation.

📄 PDF Abstract BibTeX arXiv:2407.17493

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationScheduling

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering

2025-11-16 · Zhongteng Cai, Yaxuan Wang, Yang Liu, Xueru Zhang arxiv

As synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a ``self-consuming loop" that can lead to training instability or \textit{model co…

Self-Correcting Self-Consuming Loops for Generative Model Training

2024-02-11 · Nate Gillman, Michael Freeman, Daksh Aggarwal, Chia-Hong Hsu 외

As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic…

Motion SynthesisRepresentation Learning

JeDi: Joint-Image Diffusion Models for Finetuning-Free Personalized Text-to-Image Generation

2024-07-08 · CVPR 2024 1 · Yu Zeng, Vishal M. Patel, Haochen Wang, Xun Huang 외

Personalized text-to-image generation models enable users to create images that depict their individual possessions in diverse scenes, finding applications in various domains. To achieve the personalization capability, e…

Dataset GenerationImage GenerationText to Image GenerationText-to-Image Generation

Self-Improving Diffusion Models with Synthetic Data

2024-08-29 · Sina AlEMohammad, Ahmed Imtiaz Humayun, Shruti Agarwal, John Collomosse 외

The artificial intelligence (AI) world is running out of real data for training increasingly large generative models, resulting in accelerating pressure to train on synthetic data. Unfortunately, training new generative …

FairnessImage Generation

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

2025-05-24 · Bryan Sangwoo Kim, Jeongsol Kim, Jong Chul Ye

Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collapse when asked to magnify far beyond that regime. We address this scalability bo…

Image Super-ResolutionLanguage ModelingLanguage ModellingSuper-Resolution