paper-with-me

홈 › Papers

Data-Copying in Generative Models: A Formal Framework

2023-02-25 · Robi Bhattacharjee, Sanjoy Dasgupta, Kamalika Chaudhuri

There has been some recent interest in detecting and addressing memorization of training data by deep neural networks. A formal framework for memorization in generative models, called "data-copying," was proposed by Meehan et. al. (2020). We build upon their work to show that their framework may fail to detect certain kinds of blatant memorization. Motivated by this and the theory of non-parametric methods, we provide an alternative definition of data-copying that applies more locally. We provide a method to detect data-copying, and provably show that it works with high probability when enough data is available. We also provide lower bounds that characterize the sample requirement for reliable detection.

📄 PDF Abstract BibTeX arXiv:2302.13181

Code (0)

등록된 구현이 없습니다.

Tasks

Memorization

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

A Non-Parametric Test to Detect Data-Copying in Generative Models

2020-04-12 · Casey Meehan, Kamalika Chaudhuri, Sanjoy Dasgupta

Detecting overfitting in generative models is an important challenge in machine learning. In this work, we formalize a form of overfitting that we call {\em{data-copying}} -- where the generative model memorizes and outp…

BIG-bench Machine Learning

Blameless Users in a Clean Room: Defining Copyright Protection for Generative Models

2025-06-23 · Aloni Cohen

Are there any conditions under which a generative model's outputs are guaranteed not to infringe the copyrights of its training data? This is the question of "provable copyright protection" first posed by Vyas, Kakade, a…

counterfactual

Data Plagiarism Index: Characterizing the Privacy Risk of Data-Copying in Tabular Generative Models

2024-06-18 · Joshua Ward, Chi-Hua Wang, Guang Cheng

The promise of tabular generative models is to produce realistic synthetic data that can be shared and safely used without dangerous leakage of information from the training set. In evaluating these models, a variety of …

FairnessInference AttackMembership Inference Attack

Joint Copying and Restricted Generation for Paraphrase

2016-11-28 · Ziqiang Cao, Chuwei Luo, Wenjie Li, Sujian Li

Many natural language generation tasks, such as abstractive summarization and text simplification, are paraphrase-orientated. In these tasks, copying and rewriting are two main writing modes. Most previous sequence-to-se…

Abstractive Text SummarizationDecoderInformativenessText Generation+1

Controlling the Amount of Verbatim Copying in Abstractive Summarization

2019-11-23 · Kaiqiang Song, Bingqing Wang, Zhe Feng, Liu Ren 외

An abstract must not change the meaning of the original text. A single most effective way to achieve that is to increase the amount of copying while still allowing for text abstraction. Human editors can usually exercise…

Abstractive Text SummarizationLanguage ModelingLanguage ModellingText Summarization