paper-with-me

Papers

Delphic Offline Reinforcement Learning under Nonidentifiable Hidden Confounding

2023-06-01 · Alizée Pace, Hugo Yèche, Bernhard Schölkopf, Gunnar Rätsch, Guy Tennenholtz

A prominent challenge of offline reinforcement learning (RL) is the issue of hidden confounding: unobserved variables may influence both the actions taken by the agent and the observed outcomes. Hidden confounding can compromise the validity of any causal conclusion drawn from data and presents a major obstacle to effective offline RL. In the present paper, we tackle the problem of hidden confounding in the nonidentifiable setting. We propose a definition of uncertainty due to hidden confounding bias, termed delphic uncertainty, which uses variation over world models compatible with the observations, and differentiate it from the well-known epistemic and aleatoric uncertainties. We derive a practical method for estimating the three types of uncertainties, and construct a pessimistic offline RL algorithm to account for them. Our method does not assume identifiability of the unobserved confounders, and attempts to reduce the amount of confounding bias. We demonstrate through extensive experiments and ablations the efficacy of our approach on a sepsis management benchmark, as well as on electronic health records. Our results suggest that nonidentifiable hidden confounding bias can be mitigated to improve offline RL solutions in practice.

📄 PDF Abstract BibTeX arXiv:2306.01157

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementOffline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Delphic Costs and Benefits in Web Search: A utilitarian and historical analysis

2023-08-15 · Andrei Z. Broder, Preston McAfee

We present a new framework to conceptualize and operationalize the total user experience of search, by studying the entirety of a search journey from an utilitarian point of view. Web search engines are widely perceived …

Information RetrievalMisinformation

DELPHIC: Practical DEL Planning via Possibilities (Extended Version)

2023-07-28 · Alessandro Burigana, Paolo Felli, Marco Montali

Dynamic Epistemic Logic (DEL) provides a framework for epistemic planning that is capable of representing non-deterministic actions, partial observability, higher-order knowledge and both factual and epistemic change. Th…

Transfer Learning with Partially Observable Offline Data via Causal Bounds

2023-08-07 · Xueping Gong, Wei You, Jiheng Zhang

Transfer learning has emerged as an effective approach to accelerate learning by integrating knowledge from related source agents. However, challenges arise due to data heterogeneity-such as differences in feature sets o…

Multi-Armed BanditsTransfer Learning

Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions

2026-07-28 · Zeyu Bian, Ying Zhou, Yifan Cui arxiv

Standard offline reinforcement learning (RL) algorithms typically assume that the actions in the dataset are observed without error. However, in many real-world applications, the true actions are unobserved and only nois…

Reinforcement LearningOffline RL

Learning Optimal and Sample-Efficient Decision Policies with Guarantees

2026-02-20 · Daqian Shao arxiv

The paradigm of decision-making has been revolutionised by reinforcement learning and deep learning. Although this has led to significant progress in domains such as robotics, healthcare, and finance, the use of RL in pr…

Reinforcement LearningDecision Making