paper-with-me

홈 › Papers

Top-K Ranking Deep Contextual Bandits for Information Selection Systems

2022-01-28 · Jade Freeman, Michael Rawson

In today's technology environment, information is abundant, dynamic, and heterogeneous in nature. Automated filtering and prioritization of information is based on the distinction between whether the information adds substantial value toward one's goal or not. Contextual multi-armed bandit has been widely used for learning to filter contents and prioritize according to user interest or relevance. Learn-to-Rank technique optimizes the relevance ranking on items, allowing the contents to be selected accordingly. We propose a novel approach to top-K rankings under the contextual multi-armed bandit framework. We model the stochastic reward function with a neural network to allow non-linear approximation to learn the relationship between rewards and contexts. We demonstrate the approach and evaluate the the performance of learning from the experiments using real world data sets in simulated scenarios. Empirical results show that this approach performs well under the complexity of a reward structure and high dimensional contextual features.

📄 PDF Abstract BibTeX arXiv:2201.13287

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Armed Bandits

Similar Papers 제목 키워드 기반

Model selection for contextual bandits

2019-06-03 · NeurIPS 2019 12 · Dylan J. Foster, Akshay Krishnamurthy, Haipeng Luo

We introduce the problem of model selection for contextual bandits, where a learner must adapt to the complexity of the optimal policy while balancing exploration and exploitation. Our main result is a new model selectio…

modelModel SelectionMulti-Armed Bandits

Contextual memory bandit for pro-active dialog engagement

2018-01-01 · ICLR 2018 1 · julien perez, Tomi Silander

An objective of pro-activity in dialog systems is to enhance the usability of conversational agents by enabling them to initiate conversation on their own. While dialog systems have become increasingly popular during the…

Multi-Armed Bandits

Deep Upper Confidence Bound Algorithm for Contextual Bandit Ranking of Information Selection

2021-10-08 · Michael Rawson, Jade Freeman

Contextual multi-armed bandits (CMAB) have been widely used for learning to filter and prioritize information according to a user's interest. In this work, we analyze top-K ranking under the CMAB framework where the top-…

Multi-Armed Bandits

Privacy-Preserving Multi-Party Contextual Bandits

2019-10-11 · Awni Hannun, Brian Knott, Shubho Sengupta, Laurens van der Maaten

Contextual bandits are online learners that, given an input, select an arm and receive a reward for that arm. They use the reward as a learning signal and aim to maximize the total reward over the inputs. Contextual band…

Multi-Armed BanditsPrivacy Preserving

Greedy Bandits with Sampled Context

2020-07-27 · Dom Huh

Bayesian strategies for contextual bandits have proved promising in single-state reinforcement learning tasks by modeling uncertainty using context information from the environment. In this paper, we propose Greedy Bandi…

Decision MakingMulti-Armed Banditsreinforcement-learningReinforcement Learning (RL)+1