Detecting Generative Parroting through Overfitting Masked Autoencoders
The advent of generative AI models has revolutionized digital content creation, yet it introduces challenges in maintaining copyright integrity due to generative parroting, where models mimic their training data too closely. Our research presents a novel approach to tackle this issue by employing an overfitted Masked Autoencoder (MAE) to detect such parroted samples effectively. We establish a detection threshold based on the mean loss across the training dataset, allowing for the precise identification of parroted content in modified datasets. Preliminary evaluations demonstrate promising results, suggesting our method's potential to ensure ethical use and enhance the legal compliance of generative models.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning
Recently-developed time series foundation models for scientific machine learning exhibit emergent abilities to predict physical systems. These abilities include zero-shot forecasting, in which a model forecasts future st…
In-Context LearningTime SeriesTime Series ForecastingA Non-Parametric Test to Detect Data-Copying in Generative Models
Detecting overfitting in generative models is an important challenge in machine learning. In this work, we formalize a form of overfitting that we call {\em{data-copying}} -- where the generative model memorizes and outp…
BIG-bench Machine LearningMotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
Generative masked transformers have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. Ho…
Motion SynthesisQuantizationPolly Want a Cracker: Analyzing Performance of Parroting on Paraphrase Generation Datasets
Paraphrase generation is an interesting and challenging NLP task which has numerous practical applications. In this paper, we analyze datasets commonly used for paraphrase generation research, and show that simply parrot…
Paraphrase GenerationSentenceDetecting Overfitting of Deep Generative Networks via Latent Recovery
State of the art deep generative networks are capable of producing images with such incredible realism that they can be suspected of memorizing training images. It is why it is not uncommon to include visualizations of t…
Facial InpaintingMemorizationSuper-Resolution