Incentives for Federated Learning: a Hypothesis Elicitation Approach
Federated learning provides a promising paradigm for collecting machine learning models from distributed data sources without compromising users' data privacy. The success of a credible federated learning system builds on the assumption that the decentralized and self-interested users will be willing to participate to contribute their local models in a trustworthy way. However, without proper incentives, users might simply opt out the contribution cycle, or will be mis-incentivized to contribute spam/false information. This paper introduces solutions to incentivize truthful reporting of a local, user-side machine learning model for federated learning. Our results build on the literature of information elicitation, but focus on the questions of eliciting hypothesis (rather than eliciting human predictions). We provide a scoring rule based framework that incentivizes truthful reporting of local hypotheses at a Bayesian Nash Equilibrium. We study the market implementation, accuracy as well as robustness properties of our proposed solution too. We verify the effectiveness of our methods using MNIST and CIFAR-10 datasets. Particularly we show that by reporting low-quality hypotheses, users will receive decreasing scores (rewards, or payments).
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningFederated Learningscoring ruleSimilar Papers 제목 키워드 기반
Choosing The Best Incentives for Belief Elicitation with an Application to Political Protests
Many experiments elicit subjects' prior and posterior beliefs about a random variable to assess how information affects one's own actions. However, beliefs are multi-dimensional objects, and experimenters often only elic…
Nondistortionary belief elicitation
A researcher wants to ask a decision-maker about a belief related to a choice the decision-maker made; examples include eliciting confidence or cognitive uncertainty. When can the researcher provide incentives for the de…
Multi-Observation Elicitation
We study loss functions that measure the accuracy of a prediction based on multiple data points simultaneously. To our knowledge, such loss functions have not been studied before in the area of property elicitation or in…
BIG-bench Machine LearningEliciting Categorical Data for Optimal Aggregation
Models for collecting and aggregating categorical data on crowdsourcing platforms typically fall into two broad categories: those assuming agents honest and consistent but with heterogeneous error rates, and those assumi…
Multiple-choiceSharp Results for Hypothesis Testing with Risk-Sensitive Agents
Statistical protocols are often used for decision-making involving multiple parties, each with their own incentives, private information, and ability to influence the distributional properties of the data. We study a gam…