paper-with-me

홈 › Papers

An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes

2025-09-30 · Emil Javurek, Valentyn Melnychuk, Jonas Schweisthal, Konstantin Hess, Dennis Frauen, Stefan Feuerriegel arxiv

Predicting individualized potential outcomes in sequential decision-making is central for optimizing therapeutic decisions in personalized medicine (e.g., which dosing sequence to give to a cancer patient). However, predicting potential outcomes over long horizons is notoriously difficult. Existing methods that break the curse of the horizon typically lack strong theoretical guarantees such as orthogonality and quasi-oracle efficiency. In this paper, we revisit the problem of predicting individualized potential outcomes in sequential decision-making (i.e., estimating Q-functions in Markov decision processes with observational data) through a causal inference lens. In particular, we develop a comprehensive theoretical foundation for meta-learners in this setting with a focus on beneficial theoretical properties. As a result, we yield a novel meta-learner called DRQ-learner and establish that it is: (1) doubly robust (i.e., valid inference under the misspecification of one of the nuisances), (2) Neyman-orthogonal (i.e., insensitive to first-order estimation errors in the nuisance functions), and (3) achieves quasi-oracle efficiency (i.e., behaves asymptotically as if the ground-truth nuisance functions were known). Our DRQ-learner is applicable to settings with both discrete and continuous state spaces. Further, our DRQ-learner is flexible and can be used together with arbitrary machine learning models (e.g., neural networks). We validate our theoretical results through numerical experiments, thereby showing that our meta-learner outperforms state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2509.26429

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

GDR-learners: Orthogonal Learning of Generative Models for Potential Outcomes

2025-09-26 · Valentyn Melnychuk, Stefan Feuerriegel arxiv

Various deep generative models have been proposed to estimate potential outcomes distributions from observational data. However, none of them have the favorable theoretical property of general Neyman-orthogonality and, a…

Orthogonal Learner for Estimating Heterogeneous Long-Term Treatment Effects

2026-04-01 · Haorui Ma, Dennis Frauen, Valentyn Melnychuk, Stefan Feuerriegel arxiv

Estimation of heterogeneous long-term treatment effects (HLTEs) is relevant for personalized decision-making in marketing, economics, and medicine, where short-term observational datasets are often combined with long-ter…

Orthogonal Survival Learners for Estimating Heterogeneous Treatment Effects from Time-to-Event Data

2025-05-19 · Dennis Frauen, Maresa Schröder, Konstantin Hess, Stefan Feuerriegel

Estimating heterogeneous treatment effects (HTEs) is crucial for personalized decision-making. However, this task is challenging in survival analysis, which includes time-to-event data with censored outcomes (e.g., due t…

Survival Analysis

Deep Reinforcement Learning for Adaptive Learning Systems

2020-04-17 · Xiao Li, Hanchen Xu, Jinming Zhang, Hua-hua Chang

In this paper, we formulate the adaptive learning problem---the problem of how to find an individualized learning plan (called policy) that chooses the most appropriate learning materials based on learner's latent traits…

Deep Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning+1

Doubly cross-fit debiased machine learning of heterogeneous treatment effects under principal stratification

2026-06-27 · Jiaqi Tong, Fan Li arxiv

Principal stratification provides a foundational framework for causal inference with intermediate outcomes by defining causal effects within subpopulations, yet existing work has largely focused on average effects across…

Causal Inference