paper-with-me

홈 › Papers

Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management

2024-06-08 · Huiling Meng, Ningyuan Chen, Xuefeng Gao

Intensity control is a type of continuous-time dynamic optimization problems with many important applications in Operations Research including queueing and revenue management. In this study, we adapt the reinforcement learning framework to intensity control using choice-based network revenue management as a case study, which is a classical problem in revenue management that features a large state space, a large action space and a continuous time horizon. We show that by utilizing the inherent discretization of the sample paths created by the jump points, a unique and defining feature of intensity control, one does not need to discretize the time horizon in advance, which was believed to be necessary because most reinforcement learning algorithms are designed for discrete-time problems. As a result, the computation can be facilitated and the discretization error is significantly reduced. We lay the theoretical foundation for the Monte Carlo and temporal difference learning algorithms for policy evaluation and develop policy gradient based actor critic algorithms for intensity control. Via a comprehensive numerical study, we demonstrate the benefit of our approach versus other state-of-the-art benchmarks.

📄 PDF Abstract BibTeX arXiv:2406.05358

Code (0)

등록된 구현이 없습니다.

Tasks

Managementreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

Choice-Model-Assisted Q-learning for Delayed-Feedback Revenue Management

2026-02-02 · Owen Shen, Patrick Jaillet arxiv

We study reinforcement learning for revenue management with delayed feedback, where a substantial fraction of value is determined by customer cancellations and modifications observed days after booking. We propose \emph{…

Reinforcement Learning

Revenue allocation in Formula One: a pairwise comparison approach

2019-09-25 · Dóra Gréta Petróczy, László Csató

A model is proposed to allocate Formula One World Championship prize money among the constructors. The methodology is based on pairwise comparison matrices, allows for the use of any weighting method, and makes possible …

Control Policy Correction Framework for Reinforcement Learning-based Energy Arbitrage Strategies

2024-04-29 · Seyed Soroush Karimi Madahi, Gargya Gokhale, Marie-Sophie Verwee, Bert Claessens 외

A continuous rise in the penetration of renewable energy sources, along with the use of the single imbalance pricing, provides a new opportunity for balance responsible parties to reduce their cost through energy arbitra…

Knowledge Distillationreinforcement-learningReinforcement Learning (RL)

Learning by exporting with a dose-response function

2025-05-06 · Francesca Micocci, Armando Rungi, Giovanni Cerulli

This paper investigates the causal effect of export intensity on productivity and other firm-level outcomes with a dose-response function. After positing that export intensity acts as a continuous treatment, we investiga…

counterfactual

Constant-Factor Algorithms for Revenue Management with Consecutive Stays

2025-06-01 · Ming Hu, Tongwen Wu

We study network revenue management problems motivated by applications such as railway ticket sales and hotel room bookings. Request types that require a resource for consecutive stays sequentially arrive with known arri…

Management