paper-with-me

홈 › Papers

Freshness-Aware Thompson Sampling

2014-09-29 · Djallel Bouneffouf

To follow the dynamicity of the user's content, researchers have recently started to model interactions between users and the Context-Aware Recommender Systems (CARS) as a bandit problem where the system needs to deal with exploration and exploitation dilemma. In this sense, we propose to study the freshness of the user's content in CARS through the bandit problem. We introduce in this paper an algorithm named Freshness-Aware Thompson Sampling (FA-TS) that manages the recommendation of fresh document according to the user's risk of the situation. The intensive evaluation and the detailed analysis of the experimental results reveals several important discoveries in the exploration/exploitation (exr/exp) behaviour.

📄 PDF Abstract BibTeX arXiv:1409.8572

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation SystemsThompson Sampling

Similar Papers 제목 키워드 기반

Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks

2024-10-25 · Yinglun Xu, Zhiwei Wang, Gagandeep Singh

Thompson sampling is one of the most popular learning algorithms for online sequential decision-making problems and has rich real-world applications. However, current Thompson sampling algorithms are limited by the assum…

Decision MakingSequential Decision MakingThompson Sampling

State-Aware Variational Thompson Sampling for Deep Q-Networks

2021-02-07 · Siddharth Aravindan, Wee Sun Lee

Thompson sampling is a well-known approach for balancing exploration and exploitation in reinforcement learning. It requires the posterior distribution of value-action functions to be maintained; this is generally intrac…

Thompson Sampling

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning

2026-04-18 · Weiyu Ma, Yongcheng Zeng, Yan Song, Xinyu Cui 외 arxiv

Reinforcement Learning (RL) has achieved impressive success in post-training Large Language Models (LLMs) and Vision-Language Models (VLMs), with on-policy algorithms such as PPO, GRPO, and REINFORCE++ serving as the dom…

Reinforcement Learning

A Survey of Freshness-Aware Wireless Networking with Reinforcement Learning

2025-12-24 · Alimu Alibotaiken, Suyang Wang, Oluwaseun T. Ajayi, Yu Cheng arxiv

The age of information (AoI) has become a central measure of data freshness in modern wireless systems, yet existing surveys either focus on classical AoI formulations or provide broad discussions of reinforcement learni…

Reinforcement LearningTrajectory Planning

Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits

2025-11-03 · Xuheng Li, Quanquan Gu arxiv

Variance-dependent regret bounds have received increasing attention in recent studies on contextual bandits. However, most of these studies are focused on upper confidence bound (UCB)-based bandit algorithms, while sampl…