paper-with-me

홈 › Papers

No-regret incentive-compatible online learning under exact truthfulness with non-myopic experts

2025-02-17 · Junpei Komiyama, Nishant A. Mehta, Ali Mortazavi

We study an online forecasting setting in which, over $T$ rounds, $N$ strategic experts each report a forecast to a mechanism, the mechanism selects one forecast, and then the outcome is revealed. In any given round, each expert has a belief about the outcome, but the expert wishes to select its report so as to maximize the total number of times it is selected. The goal of the mechanism is to obtain low belief regret: the difference between its cumulative loss (based on its selected forecasts) and the cumulative loss of the best expert in hindsight (as measured by the experts' beliefs). We consider exactly truthful mechanisms for non-myopic experts, meaning that truthfully reporting its belief strictly maximizes the expert's subjective probability of being selected in any future round. Even in the full-information setting, it is an open problem to obtain the first no-regret exactly truthful mechanism in this setting. We develop the first no-regret mechanism for this setting via an online extension of the Independent-Event Lotteries Forecasting Competition Mechanism (I-ELF). By viewing this online I-ELF as a novel instance of Follow the Perturbed Leader (FPL) with noise based on random walks with loss-dependent perturbations, we obtain $\tilde{O}(\sqrt{T N})$ regret. Our results are fueled by new tail bounds for Poisson binomial random variables that we develop. We extend our results to the bandit setting, where we give an exactly truthful mechanism obtaining $\tilde{O}(T^{2/3} N^{1/3})$ regret; this is the first no-regret result even among approximately truthful mechanisms.

📄 PDF Abstract BibTeX arXiv:2502.11483

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

No-Regret and Incentive-Compatible Online Learning

2020-02-20 · ICML 2020 1 · Rupert Freeman, David M. Pennock, Chara Podimata, Jennifer Wortman Vaughan

We study online learning settings in which experts act strategically to maximize their influence on the learning algorithm's predictions by potentially misreporting their beliefs about a sequence of binary events. Our go…

scoring rule

Incentive-compatible Bandits: Importance Weighting No More

2024-05-10 · Julian Zimmert, Teodor V. Marinov

We study the problem of incentive-compatible online learning with bandit feedback. In this class of problems, the experts are self-interested agents who might misrepresent their preferences with the goal of being selecte…

On the price of exact truthfulness in incentive-compatible online learning with bandit feedback: A regret lower bound for WSU-UX

2024-04-08 · Ali Mortazavi, Junhao Lin, Nishant A. Mehta

In one view of the classical game of prediction with expert advice with binary outcomes, in each round, each expert maintains an adversarially chosen belief and honestly reports this belief. We consider a recently introd…

Dynamic Online Recommendation for Two-Sided Market with Bayesian Incentive Compatibility

2024-06-04 · Yuantong Li, Guang Cheng, Xiaowu Dai

Recommender systems play a crucial role in internet economies by connecting users with relevant products or services. However, designing effective recommender systems faces two key challenges: (1) the exploration-exploit…

Recommendation Systems

Online Learning for Measuring Incentive Compatibility in Ad Auctions

2019-01-21 · Zhe Feng, Okke Schrijvers, Eric Sodomka

In this paper we investigate the problem of measuring end-to-end Incentive Compatibility (IC) regret given black-box access to an auction mechanism. Our goal is to 1) compute an estimate for IC regret in an auction, 2) p…