paper-with-me

홈 › Papers

Efficient Ensembles Improve Training Data Attribution

2024-05-27 · Junwei Deng, Ting-Wei Li, Shichang Zhang, Jiaqi Ma

Training data attribution (TDA) methods aim to quantify the influence of individual training data points on the model predictions, with broad applications in data-centric AI, such as mislabel detection, data selection, and copyright compensation. However, existing methods in this field, which can be categorized as retraining-based and gradient-based, have struggled with the trade-off between computational efficiency and attribution efficacy. Retraining-based methods can accurately attribute complex non-convex models but are computationally prohibitive, while gradient-based methods are efficient but often fail for non-convex models. Recent research has shown that augmenting gradient-based methods with ensembles of multiple independently trained models can achieve significantly better attribution efficacy. However, this approach remains impractical for very large-scale applications. In this work, we discover that expensive, fully independent training is unnecessary for ensembling the gradient-based methods, and we propose two efficient ensemble strategies, DROPOUT ENSEMBLE and LORA ENSEMBLE, alternative to naive independent ensemble. These strategies significantly reduce training time (up to 80%), serving time (up to 60%), and space cost (up to 80%) while maintaining similar attribution efficacy to the naive independent ensemble. Our extensive experimental results demonstrate that the proposed strategies are effective across multiple TDA methods on diverse datasets and models, including generative settings, significantly advancing the Pareto frontier of TDA methods with better computational efficiency and attribution efficacy.

📄 PDF Abstract BibTeX arXiv:2405.17293

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeComputational Efficiency

Methods 이 논문이 사용한 방법론

Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Training Data Attribution for Diffusion Models

2023-06-03 · Zheng Dai, David K Gifford

Diffusion models have become increasingly popular for synthesizing high-quality samples based on training datasets. However, given the oftentimes enormous sizes of the training datasets, it is difficult to assess how tra…

Exploring Resiliency to Natural Image Corruptions in Deep Learning using Design Diversity

2023-03-15 · Rafael Rosales, Pablo Munoz, Michael Paulitsch

In this paper, we investigate the relationship between diversity metrics, accuracy, and resiliency to natural image corruptions of Deep Learning (DL) image classifier ensembles. We investigate the potential of an attribu…

DiversityNeural Architecture SearchPrediction

Selective Ensembles for Consistent Predictions

2021-11-16 · ICLR 2022 4 · Emily Black, Klas Leino, Matt Fredrikson

Recent work has shown that models trained to the same objective, and which achieve similar measures of accuracy on consistent test data, may nonetheless behave very differently on individual predictions. This inconsisten…

Medical Diagnosis

Consistent Individualized Feature Attribution for Tree Ensembles

2018-02-12 · Scott M. Lundberg, Gabriel G. Erion, Su-In Lee

A unified approach to explain the output of any machine learning model.

BIG-bench Machine Learning

Counterfactual Shapley Additive Explanations

2021-10-27 · Emanuele Albini, Jason Long, Danial Dervovic, Daniele Magazzeni

Feature attributions are a common paradigm for model explanations due to their simplicity in assigning a single numeric score for each input feature to a model. In the actionable recourse setting, wherein the goal of the…

counterfactualCounterfactual ExplanationExplainable artificial intelligenceFeature Importance