paper-with-me

홈 › Papers

A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting

2020-11-02 · Philip Amortila, Nan Jiang, Tengyang Xie

Recently, Wang et al. (2020) showed a highly intriguing hardness result for batch reinforcement learning (RL) with linearly realizable value function and good feature coverage in the finite-horizon case. In this note we show that once adapted to the discounted setting, the construction can be simplified to a 2-state MDP with 1-dimensional features, such that learning is impossible even with an infinite amount of data.

📄 PDF Abstract BibTeX arXiv:2011.01075

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Comments on the Du-Kakade-Wang-Yang Lower Bounds

2019-11-18 · Benjamin Van Roy, Shi Dong

Du, Kakade, Wang, and Yang recently established intriguing lower bounds on sample complexity, which suggest that reinforcement learning with a misspecified representation is intractable. Another line of work, which cente…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Logistic Regression Regret: What's the Catch?

2020-02-07 · Gil I. Shamir

We address the problem of the achievable regret rates with online logistic regression. We derive lower bounds with logarithmic regret under $L_1$, $L_2$, and $L_\infty$ constraints on the parameter values. The bounds are…

regression

Breaking the $T^{2/3}$ Barrier for Sequential Calibration

2024-06-19 · Yuval Dagan, Constantinos Daskalakis, Maxwell Fishelson, Noah Golowich 외

A set of probabilistic forecasts is calibrated if each prediction of the forecaster closely approximates the empirical distribution of outcomes on the subset of timesteps where that prediction was made. We study the fund…

The Importance of Being Smoothly Calibrated

2026-03-16 · Parikshit Gopalan, Konstantinos Stavropoulos, Kunal Talwar, Pranay Tankala arxiv

Recent work has highlighted the centrality of smooth calibration [Kakade and Foster, 2008] as a robust measure of calibration error. We generalize, unify, and extend previous results on smooth calibration, both as a robu…

Optimistic Information Directed Sampling

2024-02-23 · Gergely Neu, Matteo Papini, Ludovic Schwartz

We study the problem of online learning in contextual bandit problems where the loss function is assumed to belong to a known parametric function class. We propose a new analytic framework for this setting that bridges t…

Multi-Armed Bandits