An Efficient Framework for Crediting Data Contributors of Diffusion Models
As diffusion models are deployed in real-world settings, and their performance is driven by training data, appraising the contribution of data contributors is crucial to creating incentives for sharing quality data and to implementing policies for data compensation. Depending on the use case, model performance corresponds to various global properties of the distribution learned by a diffusion model (e.g., overall aesthetic quality). Hence, here we address the problem of attributing global properties of diffusion models to data contributors. The Shapley value provides a principled approach to valuation by uniquely satisfying game-theoretic axioms of fairness. However, estimating Shapley values for diffusion models is computationally impractical because it requires retraining on many training data subsets corresponding to different contributors and rerunning inference. We introduce a method to efficiently retrain and rerun inference for Shapley value estimation, by leveraging model pruning and fine-tuning. We evaluate the utility of our method with three use cases: (i) image quality for a DDPM trained on a CIFAR dataset, (ii) demographic diversity for an LDM trained on CelebA-HQ, and (iii) aesthetic quality for a Stable Diffusion model LoRA-finetuned on Post-Impressionist artworks. Our results empirically demonstrate that our framework can identify important data contributors across models' global properties, outperforming existing attribution methods for diffusion models.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityFairnessMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is essential for fair compensation and sust…
Cancellation of principal in banking: Four radical ideas emerge from deep examination of double entry bookkeeping in banking
Four radical ideas are presented. First, that the rationale for cancellation of principal can be modified in modern banking. Second, that non-cancellation of loan principal upon payment may cure an old problem of mainten…
AdaCred: Adaptive Causal Decision Transformers with Feature Crediting
Reinforcement learning (RL) can be formulated as a sequence modeling problem, where models predict future actions based on historical state-action-reward sequences. Current approaches typically require long trajectory se…
AttributeImitation LearningOffline RLreinforcement-learning+21 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
In this paper, we propose the 1 Trillion Token Platform (1TT Platform), a novel framework designed to facilitate efficient data sharing with a transparent and equitable profit-sharing mechanism. The platform fosters coll…
Modelisation de l'incertitude et de l'imprecision de donnees de crowdsourcing : MONITOR
Crowdsourcing is defined as the outsourcing of tasks to a crowd of contributors. The crowd is very diverse on these platforms and includes malicious contributors attracted by the remuneration of tasks and not conscientio…