paper-with-me

홈 › Papers

Validation-Induced Shapley Shifts: How Validation Structure Distorts Data Valuation

2026-07-04 · Yinan Shen, Ziao Yang, Hongfu Liu arxiv

Shapley values are widely used to attribute value to training data based on their marginal contribution to performance on a validation set. Existing practice often assumes these values are stable once the training data and model are fixed. In this work, we uncover a systematic vulnerability: even modest changes to the validation set, such as introducing noises, cause directional shifts in Shapley distributions. As noises are added, Shapley values of training samples compress toward zero. We trace this to a noise-induced neighborhood reshuffling effect: perturbations alter the local rank order between validation and training samples, flattening the valuation landscape. Using the KNN-Shapley framework, we show through synthetic and real data that these shifts are consistent and reproducible. Our findings challenge the assumption of Shapley stability and reveal a new axis of fragility in data valuation. We propose normalization and boundary-aware validation strategies to mitigate these distortions and enable more robust, interpretable valuation in machine learning marketplaces.

📄 PDF Abstract BibTeX arXiv:2607.03675

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Causal Attribution of Model Performance Gaps in Medical Imaging Under Distribution Shifts

2025-12-09 · Pedro M. Gordaliza, Nataliia Molchanova, Jaume Banus, Thomas Sanchez 외 arxiv

Deep learning models for medical image segmentation suffer significant performance drops due to distribution shifts, but the causal mechanisms behind these drops remain poorly understood. We extend causal attribution fra…

Medical Image SegmentationLesion Segmentation

Crafting Distribution Shifts for Validation and Training in Single Source Domain Generalization

2024-09-29 · Nikos Efthymiadis, Giorgos Tolias, Ondřej Chum

Single-source domain generalization attempts to learn a model on a source domain and deploy it to unseen target domains. Limiting access only to source domain data imposes two key challenges - how to train a model that c…

Domain GeneralizationImage to sketch recognitionPhoto to Rest GeneralizationSingle-Source Domain Generalization

PIcsC: Partitioning-Induced Covariate Shift Correction

2026-07-28 · Behraj Khan, Behroz Mirza, Syed Ahmad Chan Bukhari, Tahir Qasim Syed arxiv

Covariate shift across training-data partitions biases model selection and parameter estimation in cross-validation, lifelong learning, and federated learning. We propose \textit{Partition-Induced Covariate-shift Correct…

Federated Learning

Evaluating and Explaining Earthquake-Induced Liquefaction Potential through Multi-Modal Transformers

2025-02-11 · Sompote Youwai, Tipok Kitkobsin, Sutat Leelataviwat, Pornkasem Jongpradist

This study presents an explainable parallel transformer architecture for soil liquefaction prediction that integrates three distinct data streams: spectral seismic encoding, soil stratigraphy tokenization, and site-speci…

Is Data Shapley Not Better than Random in Data Selection? Ask NASH

2026-05-11 · Xiao Tian, Jue Fan, Rachael Hwee Ling Sim, Zixuan Wang 외 arxiv

Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they accoun…