paper-with-me

홈 › Papers

Contextual Markov Decision Processes

2015-02-08 · Assaf Hallak, Dotan Di Castro, Shie Mannor

We consider a planning problem where the dynamics and rewards of the environment depend on a hidden static parameter referred to as the context. The objective is to learn a strategy that maximizes the accumulated reward across all contexts. The new model, called Contextual Markov Decision Process (CMDP), can model a customer's behavior when interacting with a website (the learner). The customer's behavior depends on gender, age, location, device, etc. Based on that behavior, the website objective is to determine customer characteristics, and to optimize the interaction between them. Our work focuses on one basic scenario--finite horizon with a small known number of possible contexts. We suggest a family of algorithms with provable guarantees that learn the underlying models and the latent contexts, and optimize the CMDPs. Bounds are obtained for specific naive implementations, and extensions of the framework are discussed, laying the ground for future research.

📄 PDF Abstract BibTeX arXiv:1502.02259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MATE: Solving Contextual Markov Decision Processes with Memory of Accumulated Transition Embeddings

2026-05-17 · Himchan Hwang, Hyeokju Jeong, Gene Chung, Seungyeon Kim 외 arxiv

We propose MATE, a simple yet effective memory architecture for solving Contextual Markov Decision Processes (CMDPs), a family of MDPs parameterized by an unobserved context. In CMDPs, an optimal agent can adapt online b…

Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback

2026-02-09 · Mengxiao Zhang, Yuheng Zhang, Haipeng Luo, Paul Mineiro arxiv

In this paper, we study Interaction-Grounded Learning (IGL) [Xie et al., 2021], a paradigm designed for realistic scenarios where the learner receives indirect feedback generated by an unknown mechanism, rather than expl…

Markov Decision Processes with Continuous Side Information

2017-11-15 · Aditya Modi, Nan Jiang, Satinder Singh, Ambuj Tewari

We consider a reinforcement learning (RL) setting in which the agent interacts with a sequence of episodic MDPs. At the start of each episode the agent has access to some side-information or context that determines the d…

PAC learningReinforcement LearningReinforcement Learning (RL)

On The Statistical Complexity of Offline Decision-Making

2025-01-10 · Thanh Nguyen-Tang, Raman Arora

We study the statistical complexity of offline decision-making with function approximation, establishing (near) minimax-optimal rates for stochastic contextual bandits and Markov decision processes. The performance limit…

Decision MakingMulti-Armed Bandits

A Notation for Markov Decision Processes

2015-12-30 · Philip S. Thomas, Billy Okal

This paper specifies a notation for Markov decision processes.