Overcoming Prior Misspecification in Online Learning to Rank
The recent literature on online learning to rank (LTR) has established the utility of prior knowledge to Bayesian ranking bandit algorithms. However, a major limitation of existing work is the requirement for the prior used by the algorithm to match the true prior. In this paper, we propose and analyze adaptive algorithms that address this issue and additionally extend these results to the linear and generalized linear models. We also consider scalar relevance feedback on top of click feedback. Moreover, we demonstrate the efficacy of our algorithms using both synthetic and real-world experiments.
Code (1)
Tasks
Learning-To-RankSimilar Papers 제목 키워드 기반
Fast and Robust Rank Aggregation against Model Misspecification
In rank aggregation (RA), a collection of preferences from different users are summarized into a total order under the assumption of homogeneity of users. Model misspecification in RA arises since the homogeneity assumpt…
Bayesian InferencemodelAdapting to Misspecification in Contextual Bandits with Offline Regression Oracles
Computationally efficient contextual bandits are often based on estimating a predictive model of rewards given contexts and arms using past data. However, when the reward model is not well-specified, the bandit algorithm…
Multi-Armed BanditsregressionTotal robustness in Bayesian Nonlinear Regression
Modern regression analyses are often undermined by covariate measurement error, misspecification of the regression model, and misspecification of the measurement error distribution. We present, to the best of our knowled…
Robust Low Rank Kernel Embeddings of Multivariate Distributions
Kernel embedding of distributions has led to many recent advances in machine learning. However, latent and low rank structures prevalent in real world distributions have rarely been taken into account in this setting. Fu…
BIG-bench Machine LearningDensity EstimationAdapting to Misspecification in Contextual Bandits
A major research direction in contextual bandits is to develop algorithms that are computationally efficient, yet support flexible, general-purpose function approximation. Algorithms based on modeling rewards have shown …
Multi-Armed Banditsregression