paper-with-me

Papers Thompson Sampling

“Thompson Sampling” 태그가 달린 논문 655편 · 필터 해제

Robust Policy Switching for Antifragile Reinforcement Learning for UAV Deconfliction in Adversarial Environments

2025-06-26 · Deepak Kumar Panda, Weisi Guo

The increasing automation of navigation for unmanned aerial vehicles (UAVs) has exposed them to adversarial attacks that exploit vulnerabilities in reinforcement learning (RL) through sensor manipulation. Although existi…

Reinforcement Learning (RL)Thompson Sampling

Context Attribution with Multi-Armed Bandit Optimization

2025-06-24 · Deng Pan, Keerthiram Murugesan, Nuno Moniz, Nitesh Chawla

Understanding which parts of the retrieved context contribute to a large language model's generated answer is essential for building interpretable and trustworthy generative QA systems. We propose a novel framework that …

Thompson Sampling

Adaptive Data Augmentation for Thompson Sampling

2025-06-17 · Wonyoung Kim

In linear contextual bandits, the objective is to select actions that maximize cumulative rewards, modeled as a linear function with unknown parameters. Although Thompson Sampling performs well empirically, it does not a…

Data AugmentationMulti-Armed BanditsThompson Sampling

Bayesian Optimization with Inexact Acquisition: Is Random Grid Search Sufficient?

2025-06-13 · Hwanwoo Kim, Chong Liu, Yuxin Chen

Bayesian optimization (BO) is a widely used iterative algorithm for optimizing black-box functions. Each iteration requires maximizing an acquisition function, such as the upper confidence bound (UCB) or a sample path fr…

Bayesian OptimizationThompson Sampling

Efficient kernelized bandit algorithms via exploration distributions

2025-06-11 · Bingshan Hu, Zheng He, Danica J. Sutherland

We consider a kernelized bandit problem with a compact arm set ${X} \subset \mathbb{R}^d $ and a fixed but unknown reward function $f^*$ with a finite norm in some Reproducing Kernel Hilbert Space (RKHS). We propose a cl…

Thompson Sampling

Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget

2025-06-03 · Jie Bian, Vincent Y. F. Tan

The challenge of identifying the best feasible arm within a fixed budget has attracted considerable interest in recent years. However, a notable gap remains in the literature: the exact exponential rate at which the erro…

Thompson Sampling

Stable Thompson Sampling: Valid Inference via Variance Inflation

2025-05-29 · Budhaditya Halder, Shubhayan Pan, Koulik Khamaru

We consider the problem of statistical inference when the data is collected via a Thompson Sampling-type algorithm. While Thompson Sampling (TS) is known to be both asymptotically optimal and empirically effective, its a…

Decision MakingThompson Samplingvalid

Simplifying Bayesian Optimization Via In-Context Direct Optimum Sampling

2025-05-29 · Gustavo Sutter Pessurno de Carvalho, Mohammed Abdulrahman, Hao Wang, Sriram Ganapathi Subramanian 외

The optimization of expensive black-box functions is ubiquitous in science and engineering. A common solution to this problem is Bayesian optimization (BO), which is generally comprised of two components: (i) a surrogate…

Bayesian OptimizationThompson Sampling

Thompson Sampling in Online RLHF with General Function Approximation

2025-05-29 · Songtao Feng, Jie Fu

Reinforcement learning from human feedback (RLHF) has achieved great empirical success in aligning large language models (LLMs) with human preference, and it is of great importance to study the statistical efficiency of …

Thompson Sampling

Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection

2025-05-28 · Qirun Zeng, Eric He, Richard Hoffmann, Xuchuang Wang 외

Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting their relevance to real-world systems. We p…

Thompson Sampling

Representative Action Selection for Large Action-Space Meta-Bandits

2025-05-23 · Quan Zhou, Mark Kozdoba, Shie Mannor

We study the problem of selecting a subset from a large action space shared by a family of bandits, with the goal of achieving performance nearly matching that of using the full action space. We assume that similar actio…

Thompson Sampling

Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions

2025-05-22 · Marc Brooks, Gabriel Durham, Kihyuk Hong, Ambuj Tewari

Recent advances in generative artificial intelligence (GenAI) models have enabled the generation of personalized content that adapts to up-to-date user context. While personalized decision systems are often modeled using…

Large Language ModelThompson Sampling

Scalable and Interpretable Contextual Bandits: A Literature Review and Retail Offer Prototype

2025-05-22 · Nikola Tankovic, Robert Sajina

This paper presents a concise review of Contextual Multi-Armed Bandit (CMAB) methods and introduces an experimental framework for scalable, interpretable offer selection, addressing the challenge of fast-changing offers.…

Feature EngineeringLarge Language ModelMulti-Armed BanditsThompson Sampling+1

Deconfounded Warm-Start Thompson Sampling with Applications to Precision Medicine

2025-05-22 · Prateek Jaiswal, Esmaeil Keyvanshokooh, Junyu Cao

Randomized clinical trials often require large patient cohorts before drawing definitive conclusions, yet abundant observational data from parallel studies remains underutilized due to confounding and hidden biases. To b…

Thompson Sampling

Steering Generative Models with Experimental Data for Protein Fitness Optimization

2025-05-21 · Jason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo 외

Protein fitness optimization involves finding a protein sequence that maximizes desired quantitative properties in a combinatorially large design space of possible sequences. Recent developments in steering protein gener…

Bayesian OptimizationThompson Sampling

In-Domain African Languages Translation Using LLMs and Multi-armed Bandits

2025-05-21 · Pratik Rakesh Singh, Kritarth Prasad, Mohammadi Zaki, Pankaj Wasnik

Neural Machine Translation (NMT) systems face significant challenges when working with low-resource languages, particularly in domain adaptation tasks. These difficulties arise due to limited training data and suboptimal…

Domain AdaptationMachine TranslationModel SelectionMulti-Armed Bandits+3

Dynamic Decision-Making under Model Misspecification

2025-05-20 · Xinyu Dai

In this study, I investigate the dynamic decision problem with a finite parameter space when the functional form of conditional expected rewards is misspecified. Traditional algorithms, such as Thompson Sampling, guarant…

Decision MakingmodelThompson Sampling

Addressing Missing Data Issue for Diffusion-based Recommendation

2025-05-18 · Wenyu Mao, Zhengyi Yang, Jiancan Wu, Haozhe Liu 외

Diffusion models have shown significant potential in generating oracle items that best match user preference with guidance from user historical interaction sequences. However, the quality of guidance is often compromised…

DenoisingThompson Sampling

Thompson Sampling-like Algorithms for Stochastic Rising Bandits

2025-05-17 · Marco Fiandri, Alberto Maria Metelli, Francesco Trovò

Stochastic rising rested bandit (SRRB) is a setting where the arms' expected rewards increase as they are pulled. It models scenarios in which the performances of the different options grow as an effect of an underlying …

Model SelectionThompson Sampling

Leveraging Offline Data from Similar Systems for Online Linear Quadratic Control

2025-05-14 · Shivam Bajaj, Prateek Jaiswal, Vijay Gupta

``Sim2real gap", in which the system learned in simulations is not the exact representation of the real system, can lead to loss of stability and performance when controllers learned using data from the simulated system …

Thompson Sampling
1–20 / 655 다음 →