paper-with-me

홈 › Papers

Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits

2024-05-06 · Jiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon, Ruoxi Jia

Data Shapley provides a principled approach to data valuation and plays a crucial role in data-centric machine learning (ML) research. Data selection is considered a standard application of Data Shapley. However, its data selection performance has shown to be inconsistent across settings in the literature. This study aims to deepen our understanding of this phenomenon. We introduce a hypothesis testing framework and show that Data Shapley's performance can be no better than random selection without specific constraints on utility functions. We identify a class of utility functions, monotonically transformed modular functions, within which Data Shapley optimally selects data. Based on this insight, we propose a heuristic for predicting Data Shapley's effectiveness in data selection tasks. Our experiments corroborate these findings, adding new insights into when Data Shapley may or may not succeed.

📄 PDF Abstract BibTeX arXiv:2405.03875

Code (0)

등록된 구현이 없습니다.

Tasks

Data Valuation

Similar Papers 제목 키워드 기반

LLpowershap: Logistic Loss-based Automated Shapley Values Feature Selection Method

2024-01-23 · Iqbal Madakkatel, Elina Hyppönen

Shapley values have been used extensively in machine learning, not only to explain black box machine learning models, but among other tasks, also to conduct model debugging, sensitivity and fairness analyses and to selec…

BenchmarkingFairnessfeature selection

DemoShapley: Valuation of Demonstrations for In-Context Learning

2024-10-10 · Shan Xie, Man Luo, Chadly Daniel Stern, Mengnan Du 외

Large language models (LLMs) leveraging in-context learning (ICL) have set new benchmarks in few-shot learning across various tasks without needing task-specific fine-tuning. However, extensive research has demonstrated …

FairnessFew-Shot LearningIn-Context Learning

Shapley values for feature selection: The good, the bad, and the axioms

2021-02-22 · Daniel Fryer, Inga Strümke, Hien Nguyen

The Shapley value has become popular in the Explainable AI (XAI) literature, thanks, to a large extent, to a solid theoretical foundation, including four "favourable and fair" axioms for attribution in transferable utili…

Explainable Artificial Intelligence (XAI)feature selection

Is Data Shapley Not Better than Random in Data Selection? Ask NASH

2026-05-11 · Xiao Tian, Jue Fan, Rachael Hwee Ling Sim, Zixuan Wang 외 arxiv

Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they accoun…

Data Selection for Fine-tuning Large Language Models Using Transferred Shapley Values

2023-06-16 · Stephanie Schoch, Ritwick Mishra, Yangfeng Ji

Although Shapley values have been shown to be highly effective for identifying harmful training instances, dataset size and model complexity constraints limit the ability to apply Shapley-based data valuation to fine-tun…

Data ValuationLanguage ModelingLanguage ModellingNatural Language Understanding