paper-with-me

홈 › Papers

On the Impact of the Utility in Semivalue-based Data Valuation

2025-02-10 · Mélissa Tamine, Benjamin Heymann, Patrick Loiseau, Maxime Vono

Semivalue-based data valuation uses cooperative-game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the practitioner's choice of utility, raising the question: How robust is semivalue-based data valuation to changes in the utility? This issue is critical when the utility is set as a trade-off between several criteria and when practitioners must select among multiple equally valid utilities. We address it by introducing the notion of a dataset's spatial signature: given a semivalue, we embed each data point into a lower-dimensional space where any utility becomes a linear functional, making the data valuation framework amenable to a simpler geometric picture. Building on this, we propose a practical methodology centered on an explicit robustness metric that informs practitioners whether and by how much their data valuation results will shift as the utility changes. We validate this approach across diverse datasets and semivalues, demonstrating strong agreement with rank-correlation analyses and offering analytical insight into how choosing a semivalue can amplify or diminish robustness.

📄 PDF Abstract BibTeX arXiv:2502.06574

Code (0)

등록된 구현이 없습니다.

Tasks

Data Valuationvalid

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Semivalue-based data valuation is arbitrary and gameable

2025-06-14 · Hannah Diehl, Ashia C. Wilson

The game-theoretic notion of the semivalue offers a popular framework for credit attribution and data valuation in machine learning. Semivalues have been proposed for a variety of high-stakes decisions involving data, su…

Data Valuation

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

2026-05-12 · Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Xinyi Xu 외 arxiv

Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its d…

Data Banzhaf: A Robust Data Valuation Framework for Machine Learning

2022-05-30 · Jiachen T. Wang, Ruoxi Jia

Data valuation has wide use cases in machine learning, including improving data quality and creating economic incentives for data sharing. This paper studies the robustness of data valuation to noisy model performance sc…

Data Valuation

Is Data Shapley Not Better than Random in Data Selection? Ask NASH

2026-05-11 · Xiao Tian, Jue Fan, Rachael Hwee Ling Sim, Zixuan Wang 외 arxiv

Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they accoun…

On the Complexity of the Inverse Semivalue Problem for Weighted Voting Games

2018-12-31 · Ilias Diakonikolas, Chrystalla Pavlou

Weighted voting games are a family of cooperative games, typically used to model voting situations where a number of agents (players) vote against or for a proposal. In such games, a proposal is accepted if an appropriat…