Recommendation System-based Upper Confidence Bound for Online Advertising
In this paper, the method UCB-RS, which resorts to recommendation system (RS) for enhancing the upper-confidence bound algorithm UCB, is presented. The proposed method is used for dealing with non-stationary and large-state spaces multi-armed bandit problems. The proposed method has been targeted to the problem of the product recommendation in the online advertising. Through extensive testing with RecoGym, an OpenAI Gym-based reinforcement learning environment for the product recommendation in online advertising, the proposed method outperforms the widespread reinforcement learning schemes such as $\epsilon$-Greedy, Upper Confidence (UCB1) and Exponential Weights for Exploration and Exploitation (EXP3).
Code (0)
등록된 구현이 없습니다.
Tasks
OpenAI GymProduct Recommendationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Knowledge Infused Policy Gradients with Upper Confidence Bound for Relational Bandits
Contextual Bandits find important use cases in various real-life scenarios such as online advertising, recommendation systems, healthcare, etc. However, most of the algorithms use flat feature vectors to represent contex…
DescriptiveMulti-Armed BanditsMusic RecommendationRecommendation SystemsCombinatorial Rising Bandit
Combinatorial online learning is a fundamental task to decide the optimal combination of base arms in sequential interactions with systems providing uncertain rewards, which is applicable to diverse domains such as robot…
Deep Reinforcement LearningRecommendation SystemsMulti-User Contextual Cascading Bandits for Personalized Recommendation
We introduce a Multi-User Contextual Cascading Bandit model, a new combinatorial bandit framework that captures realistic online advertising scenarios where multiple users interact with sequentially displayed items simul…
Epinet for Content Cold Start
The exploding popularity of online content and its user base poses an evermore challenging matching problem for modern recommendation systems. Unlike other frontiers of machine learning such as natural language, recommen…
Recommendation SystemsThompson SamplingUncertainty QuantificationNeural Contextual Bandits Under Delayed Feedback Constraints
This paper presents a new algorithm for neural contextual bandits (CBs) that addresses the challenge of delayed reward feedback, where the reward for a chosen action is revealed after a random, unknown delay. This scenar…
Multi-Armed BanditsRecommendation SystemsThompson Sampling