paper-with-me

Papers

Exploiting Action Impact Regularity and Exogenous State Variables for Offline Reinforcement Learning

2021-11-15 · Vincent Liu, James R. Wright, Martha White

Offline reinforcement learning -- learning a policy from a batch of data -- is known to be hard for general MDPs. These results motivate the need to look at specific classes of MDPs where offline reinforcement learning might be feasible. In this work, we explore a restricted class of MDPs to obtain guarantees for offline reinforcement learning. The key property, which we call Action Impact Regularity (AIR), is that actions primarily impact a part of the state (an endogenous component) and have limited impact on the remaining part of the state (an exogenous component). AIR is a strong assumption, but it nonetheless holds in a number of real-world domains including financial markets. We discuss algorithms that exploit the AIR property, and provide a theoretical analysis for an algorithm based on Fitted-Q Iteration. Finally, we demonstrate that the algorithm outperforms existing offline reinforcement learning algorithms across different data collection policies in simulated and real world environments where the regularity holds.

📄 PDF Abstract BibTeX arXiv:2111.08066

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning

2024-09-22 · Jia Wan, Sean R. Sinclair, Devavrat Shah, Martin J. Wainwright

We study Exo-MDPs, a structured class of Markov Decision Processes (MDPs) where the state space is partitioned into exogenous and endogenous components. Exogenous states evolve stochastically, independent of the agent's …

reinforcement-learningReinforcement Learning

Learning in Markov Decision Processes with Exogenous Dynamics

2026-03-03 · Davide Maran, Davide Salaorni, Marcello Restelli arxiv

Reinforcement learning algorithms are typically designed for generic Markov Decision Processes (MDPs), where any state-action pair can lead to an arbitrary transition distribution. In many practical systems, however, onl…

Reinforcement Learning

Shaping Social Activity by Incentivizing Users

2014-12-01 · NeurIPS 2014 12 · Mehrdad Farajtabar, Nan Du, Manuel Gomez Rodriguez, Isabel Valera 외

Events in an online social network can be categorized roughly into endogenous events, where users just respond to the actions of their neighbors within the network, or exogenous events, where users take actions due to dr…

Spatiotemporal forecasting of vertical track alignment with exogenous factors

2022-11-07 · Katsuya Kosukegawa, Yasukuni Mori, Hiroki Suyari, Kazuhiko Kawamoto

To ensure the safety of railroad operations, it is important to monitor and forecast track geometry irregularities. A higher safety requires forecasting with higher spatiotemporal frequencies, which in turn requires capt…

The Random Walk behind Volatility Clustering

2016-12-29

Financial price changes obey two universal properties: they follow a power law and they tend to be clustered in time. The second regularity, known as volatility clustering, entails some predictability in the price change…

Clustering