paper-with-me

Papers

Quotient Semivalues for False-Name-Resistant Data Attribution

2026-05-08 · Florian A. D. Burnat, Brittany I. Davidson arxiv

Data valuation methods allocate payments and audit training data's contribution to machine-learning pipelines; however, they often assume passive contributors. In reality, contributors can split datasets across pseudonymous identities, duplicate high-value examples, create near-duplicates, or launder synthetic variants to inflate their share. We formalize this as false-name manipulation in ML data attribution. Our main construction is the quotient semivalue mechanism: compute Shapley-, Banzhaf-, or Beta-style values over evidence-backed attribution clusters instead of raw identities, using a canonical-representative operator to absorb within-cluster duplication. We prove an impossibility: on a fixed monotone data-value game, exact Shapley-fair attribution over reported identities is incompatible with unrestricted false-name-proofness, even on binary-valued instances, and characterize the split-gain of a general semivalue on a unanimity counter-example. The mechanism is exactly false-name-proof under two structural conditions: false-name-neutral within-cluster allocation and quotient-stable manipulations. Under imperfect provenance, when these conditions hold approximately, manipulation gain and fairness loss are bounded by three measurable quantities: escaped-cluster mass, value-estimation error, and clustering distance. We instantiate the mechanisms in DataMarket-Gym, a benchmark for attribution under strategic provider attacks. On synthetic classification tasks, quotient semivalues with example-level evidence reduce manipulation gain on duplicate and near-duplicate Sybil attacks from $1.74$ under baseline Shapley to $0.96$, near the honest level. The cosine-threshold and (false-merge, false-split) rate sweeps trace the corresponding fairness--Sybil frontier.

📄 PDF Abstract BibTeX arXiv:2605.07663

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients

2026-06-15 · Yan Jiao, Pin-Han Ho, Limei Peng arxiv

Learned representations in intelligent sensing systems are often evaluated by reconstruction fidelity or downstream prediction accuracy, but these criteria do not specify which latent distinctions are justified by the se…

Representation Learning

On the Complexity of the Inverse Semivalue Problem for Weighted Voting Games

2018-12-31 · Ilias Diakonikolas, Chrystalla Pavlou

Weighted voting games are a family of cooperative games, typically used to model voting situations where a number of agents (players) vote against or for a proposal. In such games, a proposal is accepted if an appropriat…

Quotient Based Multiresolution Image Fusion of Thermal and Visual Images Using Daubechies Wavelet Transform for Human Face Recognition

2010-07-05 · Mrinal Kanti Bhowmik, Debotosh Bhattacharjee, Mita Nasipuri, Dipak Kumar Basu 외

This paper investigates the multiresolution level-1 and level-2 Quotient based Fusion of thermal and visual images. In the proposed system, the method-1 namely "Decompose then Quotient Fuse Level-1" and the method-2 name…

Dimensionality ReductionFace Recognition

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

2026-05-12 · Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Xinyi Xu 외 arxiv

Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its d…

Isometric Quotient Variational Auto-Encoders for Structure-Preserving Representation Learning

2023-09-21 · NeurIPS 2023 11

We study structure-preserving low-dimensional representation of a data manifold embedded in a high-dimensional observation space based on variational auto-encoders (VAEs). We approach this by decomposing the data manifol…