Approximate optimality and the risk/reward tradeoff in a class of bandit problems
This paper studies a sequential decision problem where payoff distributions are known and where the riskiness of payoffs matters. Equivalently, it studies sequential choice from a repeated set of independent lotteries. The decision-maker is assumed to pursue strategies that are approximately optimal for large horizons. By exploiting the tractability afforded by asymptotics, conditions are derived characterizing when specialization in one action or lottery throughout is asymptotically optimal and when optimality requires intertemporal diversification. The key is the constancy or variability of risk attitude. The main technical tool is a new central limit theorem.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Explicit-risk-aware Path Planning with Reward Maximization
This paper develops a path planner that minimizes risk (e.g. motion execution) while maximizing accumulated reward (e.g., quality of sensor viewpoint) motivated by visual assistance or tracking scenarios in unstructured …
Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret
We study risk-sensitive reinforcement learning in episodic Markov decision processes with unknown transition kernels, where the goal is to optimize the total reward under the risk measure of exponential utility. We propo…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Exploration Potential
We introduce exploration potential, a quantity that measures how much a reinforcement learning agent has explored its environment class. In contrast to information gain, exploration potential takes the problem's reward s…
Multi-Armed Banditsreinforcement-learningReinforcement LearningReinforcement Learning (RL)Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization
Safe reinforcement learning (Safe RL) refers to a class of techniques that aim to prevent RL algorithms from violating constraints in the process of decision-making and exploration during trial and error. In this paper, …
Decision Makingreinforcement-learningReinforcement LearningSafe Reinforcement LearningReward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy. This practice is suboptimal for maximizi…