Deep Exploration for Recommendation Systems
Modern recommendation systems ought to benefit by probing for and learning from delayed feedback. Research has tended to focus on learning from a user's response to a single recommendation. Such work, which leverages methods of supervised and bandit learning, forgoes learning from the user's subsequent behavior. Where past work has aimed to learn from subsequent behavior, there has been a lack of effective methods for probing to elicit informative delayed feedback. Effective exploration through probing for delayed feedback becomes particularly challenging when rewards are sparse. To address this, we develop deep exploration methods for recommendation systems. In particular, we formulate recommendation as a sequential decision problem and demonstrate benefits of deep exploration over single-step exploration. Our experiments are carried out with high-fidelity industrial-grade simulators and establish large improvements over existing algorithms.
Code (0)
등록된 구현이 없습니다.
Tasks
Recommendation SystemsThompson SamplingSimilar Papers 제목 키워드 기반
Fiduciary Bandits
Recommendation systems often face exploration-exploitation tradeoffs: the system can only learn about the desirability of new options by recommending them to some user. Such systems can thus be modeled as multi-armed ban…
Recommendation SystemsRecommender for Its Purpose: Repeat and Exploration in Food Delivery Recommendations
Recommender systems have been widely used for various scenarios, such as e-commerce, news, and music, providing online contents to help and enrich users' daily life. Different scenarios hold distinct and unique character…
Recommendation SystemsIncentivizing Exploration with Selective Data Disclosure
We propose and design recommendation systems that incentivize efficient exploration. Agents arrive sequentially, choose actions and receive rewards, drawn from fixed but unknown action-specific distributions. The recomme…
Efficient ExplorationRecommendation SystemsDeep density networks and uncertainty in recommender systems
Building robust online content recommendation systems requires learning complex interactions between user preferences and content features. The field has evolved rapidly in recent years from traditional multi-arm bandit …
Collaborative FilteringEfficient ExplorationRecommendation SystemsUser Feedback Alignment for LLM-powered Exploration in Large-scale Recommendation Systems
Exploration, the act of broadening user experiences beyond their established preferences, is challenging in large-scale recommendation systems due to feedback loops and limited signals on user exploration patterns. Large…
DiversityRecommendation SystemsWorld Knowledge