paper-with-me

홈 › Papers

The Role of Coverage in Online Reinforcement Learning

2022-10-09 · Tengyang Xie, Dylan J. Foster, Yu Bai, Nan Jiang, Sham M. Kakade

Coverage conditions -- which assert that the data logging distribution adequately covers the state space -- play a fundamental role in determining the sample complexity of offline reinforcement learning. While such conditions might seem irrelevant to online reinforcement learning at first glance, we establish a new connection by showing -- somewhat surprisingly -- that the mere existence of a data distribution with good coverage can enable sample-efficient online RL. Concretely, we show that coverability -- that is, existence of a data distribution that satisfies a ubiquitous coverage condition called concentrability -- can be viewed as a structural property of the underlying MDP, and can be exploited by standard algorithms for sample-efficient exploration, even when the agent does not know said distribution. We complement this result by proving that several weaker notions of coverage, despite being sufficient for offline RL, are insufficient for online RL. We also show that existing complexity measures for online RL, including Bellman rank and Bellman-Eluder dimension, fail to optimally capture coverability, and propose a new complexity measure, the sequential extrapolation coefficient, to provide a unification.

📄 PDF Abstract BibTeX arXiv:2210.04157

Code (0)

등록된 구현이 없습니다.

Tasks

Efficient ExplorationOffline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

2024-11-07 · Heyang Zhao, Chenlu Ye, Quanquan Gu, Tong Zhang

Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning from human feedback (RLHF), which force…

Multi-Armed BanditsReinforcement Learning (RL)

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage

2026-02-12 · Haolin Liu, Braham Snyder, Chen-Yu Wei arxiv

We study offline reinforcement learning under $Q^\star$-approximation and partial coverage, a setting that motivates practical algorithms such as Conservative $Q$-Learning (CQL; Kumar et al., 2020) but has received limit…

Reinforcement LearningOffline RL

Active Coverage for PAC Reinforcement Learning

2023-06-23 · Aymen Al-Marjani, Andrea Tirinzoni, Emilie Kaufmann

Collecting and leveraging data with good coverage properties plays a crucial role in different aspects of reinforcement learning (RL), including reward-free exploration and offline learning. However, the notion of "good …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

What can online reinforcement learning with function approximation benefit from general coverage conditions?

2023-04-25 · Fanghui Liu, Luca Viano, Volkan Cevher

In online reinforcement learning (RL), instead of employing standard structural assumptions on Markov decision processes (MDPs), using a certain coverage condition (original from offline RL) is enough to ensure sample-ef…

Offline RLReinforcement Learning (RL)

RALLY: Role-Adaptive LLM-Driven Yoked Navigation for Agentic UAV Swarms

2025-07-02 · Ziyao Wang, Rongpeng Li, Sizhao Li, Yuming Xiang 외 arxiv

Intelligent control of Unmanned Aerial Vehicles (UAVs) swarms has emerged as a critical research focus, and it typically requires the swarm to navigate effectively while avoiding obstacles and achieving continuous covera…

Multi-agent Reinforcement LearningSemantic Communication