paper-with-me

홈 › Papers

Random Walk Approach to Regret Minimization

2010-12-01 · NeurIPS 2010 12 · Hariharan Narayanan, Alexander Rakhlin

We propose a computationally efficient random walk on a convex body which rapidly mixes to a time-varying Gibbs distribution. In the setting of online convex optimization and repeated games, the algorithm yields low regret and presents a novel efficient method for implementing mixture forecasting strategies.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stochastic Gradient Descent on a Tree: an Adaptive and Robust Approach to Stochastic Convex Optimization

2019-01-17 · Sattar Vakili, Sudeep Salgia, Qing Zhao

Online minimization of an unknown convex function over the interval $[0,1]$ is considered under first-order stochastic bandit feedback, which returns a random realization of the gradient of the function at each query poi…

Fitting mixed logit random regret minimization models using maximum simulated likelihood

2023-01-03 · Ziyue Zhu, Álvaro A. Gutiérrez-Vargas, Martina Vandebroek

This article describes the mixrandregret command, which extends the randregret command introduced in Guti\'errez-Vargas et al. (2021, The Stata Journal 21: 626-658) incorporating random coefficients for Random Regret Min…

Numerical Integration

Optimal Non-Asymptotic Lower Bound on the Minimax Regret of Learning with Expert Advice

2015-11-06 · Francesco Orabona, David Pal

We prove non-asymptotic lower bounds on the expectation of the maximum of $d$ independent Gaussian variables and the expectation of the maximum of $d$ independent symmetric random walks. Both lower bounds recover the opt…

Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration

2020-10-23 · Priyank Agrawal, Jinglin Chen, Nan Jiang

This paper studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (T…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Thompson Sampling

Towards Optimal Algorithms for Prediction with Expert Advice

2014-09-10 · Nick Gravin, Yuval Peres, Balasubramanian Sivan

We study the classical problem of prediction with expert advice in the adversarial setting with a geometric stopping time. In 1965, Cover gave the optimal algorithm for the case of 2 experts. In this paper, we design the…

PredictionThompson Sampling