Simulated Contextual Bandits for Personalization Tasks from Recommendation Datasets
We propose a method for generating simulated contextual bandit environments for personalization tasks from recommendation datasets like MovieLens, Netflix, Last.fm, Million Song, etc. This allows for personalization environments to be developed based on real-life data to reflect the nuanced nature of real-world user interactions. The obtained environments can be used to develop methods for solving personalization tasks, algorithm benchmarking, model simulation, and more. We demonstrate our approach with numerical examples on MovieLens and IMDb datasets.
Code (1)
Tasks
BenchmarkingMulti-Armed BanditsSimilar Papers 제목 키워드 기반
Hierarchical Contextual Uplift Bandits for Catalog Personalization
Contextual Bandit (CB) algorithms are widely adopted for personalized recommendations but often struggle in dynamic environments typical of fantasy sports, where rapid changes in user behavior and dramatic shifts in rewa…
AutoML for Contextual Bandits
Contextual Bandits is one of the widely popular techniques used in applications such as personalization, recommendation systems, mobile health, causal marketing etc . As a dynamic approach, it can be more efficient than …
AutoMLFeature EngineeringMarketingMeta-Learning+2Carousel Personalization in Music Streaming Apps with Contextual Bandits
Media services providers, such as music streaming platforms, frequently leverage swipeable carousels to recommend personalized content to their users. However, selecting the most relevant items (albums, artists, playlist…
Multi-Armed BanditsPrivacy-Preserving Bandits
Contextual bandit algorithms~(CBAs) often rely on personal data to provide recommendations. Centralized CBA agents utilize potentially sensitive data from recent interactions to provide personalization to end-users. Keep…
Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONPrivacy PreservingSelectively Contextual Bandits
Contextual bandits are widely used in industrial personalization systems. These online learning frameworks learn a treatment assignment policy in the presence of treatment effects that vary with the observed contextual f…
Multi-Armed Bandits