paper-with-me

홈 › Papers

Self-Consuming Generative Models with Adversarially Curated Data

2025-05-14 · Xiukun Wei, Xueru Zhang

Recent advances in generative models have made it increasingly difficult to distinguish real data from model-generated synthetic data. Using synthetic data for successive training of future model generations creates "self-consuming loops", which may lead to model collapse or training instability. Furthermore, synthetic data is often subject to human feedback and curated by users based on their preferences. Ferbach et al. (2024) recently showed that when data is curated according to user preferences, the self-consuming retraining loop drives the model to converge toward a distribution that optimizes those preferences. However, in practice, data curation is often noisy or adversarially manipulated. For example, competing platforms may recruit malicious users to adversarially curate data and disrupt rival models. In this paper, we study how generative models evolve under self-consuming retraining loops with noisy and adversarially curated data. We theoretically analyze the impact of such noisy data curation on generative models and identify conditions for the robustness of the retraining process. Building on this analysis, we design attack algorithms for competitive adversarial scenarios, where a platform with a limited budget employs malicious users to misalign a rival's model from actual user preferences. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed algorithms.

📄 PDF Abstract BibTeX arXiv:2505.09768

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences

2024-06-12 · Damien Ferbach, Quentin Bertrand, Avishek Joey Bose, Gauthier Gidel

The rapid progress in generative models has resulted in impressive leaps in generation quality, blurring the lines between synthetic and real data. Web-scale datasets are now prone to the inevitable contamination by synt…

Self-Correcting Self-Consuming Loops for Generative Model Training

2024-02-11 · Nate Gillman, Michael Freeman, Daksh Aggarwal, Chia-Hong Hsu 외

As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic…

Motion SynthesisRepresentation Learning

Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation

2025-11-12 · Hongru Zhao, Jinwen Fu, Tuan Pham arxiv

Self-consuming generative models have received significant attention over the last few years. In this paper, we study a self-consuming generative model with heterogeneous preferences that is a generalization of the model…

Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach

2024-05-24 · Huy V. Vo, Vasil Khalidov, Timothée Darcet, Théo Moutakanni 외

Self-supervised features are the cornerstone of modern machine learning systems. They are typically pre-trained on data collections whose construction and curation typically require extensive human effort. This manual pr…

ClusteringSelf-Supervised Learning

Self-Consuming Generative Models Go MAD

2023-07-04 · Sina AlEMohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 외

Seismic advances in generative AI algorithms for imagery, text, and other data types has led to the temptation to use synthetic data to train next-generation models. Repeating this process creates an autophagous (self-co…

Diversity