paper-with-me

Papers

Recommendation System-based Upper Confidence Bound for Online Advertising

2019-09-09 · Nhan Nguyen-Thanh, Dana Marinca, Kinda Khawam, David Rohde, Flavian vasile, Elena Simona Lohan, Steven Martin, Dominique Quadri

In this paper, the method UCB-RS, which resorts to recommendation system (RS) for enhancing the upper-confidence bound algorithm UCB, is presented. The proposed method is used for dealing with non-stationary and large-state spaces multi-armed bandit problems. The proposed method has been targeted to the problem of the product recommendation in the online advertising. Through extensive testing with RecoGym, an OpenAI Gym-based reinforcement learning environment for the product recommendation in online advertising, the proposed method outperforms the widespread reinforcement learning schemes such as $\epsilon$-Greedy, Upper Confidence (UCB1) and Exponential Weights for Exploration and Exploitation (EXP3).

📄 PDF Abstract BibTeX arXiv:1909.04190

Code (0)

등록된 구현이 없습니다.

Tasks

OpenAI GymProduct Recommendationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Knowledge Infused Policy Gradients with Upper Confidence Bound for Relational Bandits

2021-06-25 · Kaushik Roy, Qi Zhang, Manas Gaur, Amit Sheth

Contextual Bandits find important use cases in various real-life scenarios such as online advertising, recommendation systems, healthcare, etc. However, most of the algorithms use flat feature vectors to represent contex…

DescriptiveMulti-Armed BanditsMusic RecommendationRecommendation Systems

Combinatorial Rising Bandit

2024-12-01 · Seockbean Song, Youngsik Yoon, Siwei Wang, Wei Chen 외

Combinatorial online learning is a fundamental task to decide the optimal combination of base arms in sequential interactions with systems providing uncertain rewards, which is applicable to diverse domains such as robot…

Deep Reinforcement LearningRecommendation Systems

Multi-User Contextual Cascading Bandits for Personalized Recommendation

2025-08-19 · Jiho Park, Huiwen Jia arxiv

We introduce a Multi-User Contextual Cascading Bandit model, a new combinatorial bandit framework that captures realistic online advertising scenarios where multiple users interact with sequentially displayed items simul…

Epinet for Content Cold Start

2024-11-20 · Hong Jun Jeon, Songbin Liu, Yuantong Li, Jie Lyu 외

The exploding popularity of online content and its user base poses an evermore challenging matching problem for modern recommendation systems. Unlike other frontiers of machine learning such as natural language, recommen…

Recommendation SystemsThompson SamplingUncertainty Quantification

Neural Contextual Bandits Under Delayed Feedback Constraints

2025-04-16 · Mohammadali Moghimi, Sharu Theresa Jose, Shana Moothedath

This paper presents a new algorithm for neural contextual bandits (CBs) that addresses the challenge of delayed reward feedback, where the reward for a chosen action is revealed after a random, unknown delay. This scenar…

Multi-Armed BanditsRecommendation SystemsThompson Sampling