paper-with-me

홈 › Papers

Semivalue-based data valuation is arbitrary and gameable

2025-06-14 · Hannah Diehl, Ashia C. Wilson

The game-theoretic notion of the semivalue offers a popular framework for credit attribution and data valuation in machine learning. Semivalues have been proposed for a variety of high-stakes decisions involving data, such as determining contributor compensation, acquiring data from external sources, or filtering out low-value datapoints. In these applications, semivalues depend on the specification of a utility function that maps subsets of data to a scalar score. While it is broadly agreed that this utility function arises from a composition of a learning algorithm and a performance metric, its actual instantiation involves numerous subtle modeling choices. We argue that this underspecification leads to varying degrees of arbitrariness in semivalue-based valuations. Small, but arguably reasonable changes to the utility function can induce substantial shifts in valuations across datapoints. Moreover, these valuation methodologies are also often gameable: low-cost adversarial strategies exist to exploit this ambiguity and systematically redistribute value among datapoints. Through theoretical constructions and empirical examples, we demonstrate that a bad-faith valuator can manipulate utility specifications to favor preferred datapoints, and that a good-faith valuator is left without principled guidance to justify any particular specification. These vulnerabilities raise ethical and epistemic concerns about the use of semivalues in several applications. We conclude by highlighting the burden of justification that semivalue-based approaches place on modelers and discuss important considerations for identifying appropriate uses.

📄 PDF Abstract BibTeX arXiv:2506.12619

Code (0)

등록된 구현이 없습니다.

Tasks

Data Valuation

Similar Papers 제목 키워드 기반

On the Impact of the Utility in Semivalue-based Data Valuation

2025-02-10 · Mélissa Tamine, Benjamin Heymann, Patrick Loiseau, Maxime Vono

Semivalue-based data valuation uses cooperative-game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the practitioner's choice of utilit…

Data Valuationvalid

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

2026-05-12 · Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Xinyi Xu 외 arxiv

Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its d…

Data Banzhaf: A Robust Data Valuation Framework for Machine Learning

2022-05-30 · Jiachen T. Wang, Ruoxi Jia

Data valuation has wide use cases in machine learning, including improving data quality and creating economic incentives for data sharing. This paper studies the robustness of data valuation to noisy model performance sc…

Data Valuation

On the Complexity of the Inverse Semivalue Problem for Weighted Voting Games

2018-12-31 · Ilias Diakonikolas, Chrystalla Pavlou

Weighted voting games are a family of cooperative games, typically used to model voting situations where a number of agents (players) vote against or for a proposal. In such games, a proposal is accepted if an appropriat…

Quotient Semivalues for False-Name-Resistant Data Attribution

2026-05-08 · Florian A. D. Burnat, Brittany I. Davidson arxiv

Data valuation methods allocate payments and audit training data's contribution to machine-learning pipelines; however, they often assume passive contributors. In reality, contributors can split datasets across pseudonym…