Papers Thompson Sampling
“Thompson Sampling” 태그가 달린 논문 655편 · 필터 해제
Robust Policy Switching for Antifragile Reinforcement Learning for UAV Deconfliction in Adversarial Environments
The increasing automation of navigation for unmanned aerial vehicles (UAVs) has exposed them to adversarial attacks that exploit vulnerabilities in reinforcement learning (RL) through sensor manipulation. Although existi…
Reinforcement Learning (RL)Thompson SamplingContext Attribution with Multi-Armed Bandit Optimization
Understanding which parts of the retrieved context contribute to a large language model's generated answer is essential for building interpretable and trustworthy generative QA systems. We propose a novel framework that …
Thompson SamplingAdaptive Data Augmentation for Thompson Sampling
In linear contextual bandits, the objective is to select actions that maximize cumulative rewards, modeled as a linear function with unknown parameters. Although Thompson Sampling performs well empirically, it does not a…
Data AugmentationMulti-Armed BanditsThompson SamplingBayesian Optimization with Inexact Acquisition: Is Random Grid Search Sufficient?
Bayesian optimization (BO) is a widely used iterative algorithm for optimizing black-box functions. Each iteration requires maximizing an acquisition function, such as the upper confidence bound (UCB) or a sample path fr…
Bayesian OptimizationThompson SamplingEfficient kernelized bandit algorithms via exploration distributions
We consider a kernelized bandit problem with a compact arm set ${X} \subset \mathbb{R}^d $ and a fixed but unknown reward function $f^*$ with a finite norm in some Reproducing Kernel Hilbert Space (RKHS). We propose a cl…
Thompson SamplingAsymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget
The challenge of identifying the best feasible arm within a fixed budget has attracted considerable interest in recent years. However, a notable gap remains in the literature: the exact exponential rate at which the erro…
Thompson SamplingStable Thompson Sampling: Valid Inference via Variance Inflation
We consider the problem of statistical inference when the data is collected via a Thompson Sampling-type algorithm. While Thompson Sampling (TS) is known to be both asymptotically optimal and empirically effective, its a…
Decision MakingThompson SamplingvalidSimplifying Bayesian Optimization Via In-Context Direct Optimum Sampling
The optimization of expensive black-box functions is ubiquitous in science and engineering. A common solution to this problem is Bayesian optimization (BO), which is generally comprised of two components: (i) a surrogate…
Bayesian OptimizationThompson SamplingThompson Sampling in Online RLHF with General Function Approximation
Reinforcement learning from human feedback (RLHF) has achieved great empirical success in aligning large language models (LLMs) with human preference, and it is of great importance to study the statistical efficiency of …
Thompson SamplingPractical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting their relevance to real-world systems. We p…
Thompson SamplingRepresentative Action Selection for Large Action-Space Meta-Bandits
We study the problem of selecting a subset from a large action space shared by a family of bandits, with the goal of achieving performance nearly matching that of using the full action space. We assume that similar actio…
Thompson SamplingGenerator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
Recent advances in generative artificial intelligence (GenAI) models have enabled the generation of personalized content that adapts to up-to-date user context. While personalized decision systems are often modeled using…
Large Language ModelThompson SamplingScalable and Interpretable Contextual Bandits: A Literature Review and Retail Offer Prototype
This paper presents a concise review of Contextual Multi-Armed Bandit (CMAB) methods and introduces an experimental framework for scalable, interpretable offer selection, addressing the challenge of fast-changing offers.…
Feature EngineeringLarge Language ModelMulti-Armed BanditsThompson Sampling+1Deconfounded Warm-Start Thompson Sampling with Applications to Precision Medicine
Randomized clinical trials often require large patient cohorts before drawing definitive conclusions, yet abundant observational data from parallel studies remains underutilized due to confounding and hidden biases. To b…
Thompson SamplingSteering Generative Models with Experimental Data for Protein Fitness Optimization
Protein fitness optimization involves finding a protein sequence that maximizes desired quantitative properties in a combinatorially large design space of possible sequences. Recent developments in steering protein gener…
Bayesian OptimizationThompson SamplingIn-Domain African Languages Translation Using LLMs and Multi-armed Bandits
Neural Machine Translation (NMT) systems face significant challenges when working with low-resource languages, particularly in domain adaptation tasks. These difficulties arise due to limited training data and suboptimal…
Domain AdaptationMachine TranslationModel SelectionMulti-Armed Bandits+3Dynamic Decision-Making under Model Misspecification
In this study, I investigate the dynamic decision problem with a finite parameter space when the functional form of conditional expected rewards is misspecified. Traditional algorithms, such as Thompson Sampling, guarant…
Decision MakingmodelThompson SamplingAddressing Missing Data Issue for Diffusion-based Recommendation
Diffusion models have shown significant potential in generating oracle items that best match user preference with guidance from user historical interaction sequences. However, the quality of guidance is often compromised…
DenoisingThompson SamplingThompson Sampling-like Algorithms for Stochastic Rising Bandits
Stochastic rising rested bandit (SRRB) is a setting where the arms' expected rewards increase as they are pulled. It models scenarios in which the performances of the different options grow as an effect of an underlying …
Model SelectionThompson SamplingLeveraging Offline Data from Similar Systems for Online Linear Quadratic Control
``Sim2real gap", in which the system learned in simulations is not the exact representation of the real system, can lead to loss of stability and performance when controllers learned using data from the simulated system …
Thompson Sampling