paper-with-me

홈 › Papers

What are the Statistical Limits of Offline RL with Linear Function Approximation?

2020-10-22 · Ruosong Wang, Dean P. Foster, Sham M. Kakade

Offline reinforcement learning seeks to utilize offline (observational) data to guide the learning of (causal) sequential decision making strategies. The hope is that offline reinforcement learning coupled with function approximation methods (to deal with the curse of dimensionality) can provide a means to help alleviate the excessive sample complexity burden in modern sequential decision making problems. However, the extent to which this broader approach can be effective is not well understood, where the literature largely consists of sufficient conditions. This work focuses on the basic question of what are necessary representational and distributional conditions that permit provable sample-efficient offline reinforcement learning. Perhaps surprisingly, our main result shows that even if: i) we have realizability in that the true value function of \emph{every} policy is linear in a given set of features and 2) our off-policy data has good coverage over all features (under a strong spectral condition), then any algorithm still (information-theoretically) requires a number of offline samples that is exponential in the problem horizon in order to non-trivially estimate the value of \emph{any} given policy. Our results highlight that sample-efficient offline policy evaluation is simply not possible unless significantly stronger conditions hold; such conditions include either having low distribution shift (where the offline data distribution is close to the distribution of the policy to be evaluated) or significantly stronger representational conditions (beyond realizability).

📄 PDF Abstract BibTeX arXiv:2010.11895

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingOffline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)Sequential Decision Making

Similar Papers 제목 키워드 기반

Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with Pessimism

2022-03-11 · Ming Yin, Yaqi Duan, Mengdi Wang, Yu-Xiang Wang

Offline reinforcement learning, which seeks to utilize offline/historical data to optimize sequential decision-making strategies, has gained surging prominence in recent studies. Due to the advantage that appropriate fun…

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

On The Statistical Complexity of Offline Decision-Making

2025-01-10 · Thanh Nguyen-Tang, Raman Arora

We study the statistical complexity of offline decision-making with function approximation, establishing (near) minimax-optimal rates for stochastic contextual bandits and Markov decision processes. The performance limit…

Decision MakingMulti-Armed Bandits

Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation

2021-11-21 · Dylan J. Foster, Akshay Krishnamurthy, David Simchi-Levi, Yunzong Xu

We consider the offline reinforcement learning problem, where the aim is to learn a decision making policy from logged data. Offline RL -- particularly when coupled with (value) function approximation to allow for genera…

Decision MakingOffline RLreinforcement-learningReinforcement Learning+1

What's the score? Automated Denoising Score Matching for Nonlinear Diffusions

2024-07-10 · Raghav Singhal, Mark Goldstein, Rajesh Ranganath

Reversing a diffusion process by learning its score forms the heart of diffusion-based generative modeling and for estimating properties of scientific systems. The diffusion processes that are tractable center on linear …

Denoising

A Complete Characterization of Linear Estimators for Offline Policy Evaluation

2022-03-08 · Juan C. Perdomo, Akshay Krishnamurthy, Peter Bartlett, Sham Kakade

Offline policy evaluation is a fundamental statistical problem in reinforcement learning that involves estimating the value function of some decision-making policy given data collected by a potentially different policy. …

Decision Makingreinforcement-learningReinforcement Learning (RL)