paper-with-me

홈 › Papers

Learning State-Dependent Policy Parametrizations for Dynamic Technician Routing with Rework

2024-09-03 · Jonas Stein, Florentin D Hildebrandt, Barrett W Thomas, Marlin W Ulmer

Home repair and installation services require technicians to visit customers and resolve tasks of different complexity. Technicians often have heterogeneous skills and working experiences. The geographical spread of customers makes achieving only perfect matches between technician skills and task requirements impractical. Additionally, technicians are regularly absent due to sickness. With non-perfect assignments regarding task requirement and technician skill, some tasks may remain unresolved and require a revisit and rework. Companies seek to minimize customer inconvenience due to delay. We model the problem as a sequential decision process where, over a number of service days, customers request service while heterogeneously skilled technicians are routed to serve customers in the system. Each day, our policy iteratively builds tours by adding "important" customers. The importance bases on analytical considerations and is measured by respecting routing efficiency, urgency of service, and risk of rework in an integrated fashion. We propose a state-dependent balance of these factors via reinforcement learning. A comprehensive study shows that taking a few non-perfect assignments can be quite beneficial for the overall service quality. We further demonstrate the value provided by a state-dependent parametrization.

📄 PDF Abstract BibTeX arXiv:2409.01815

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Multiscale Neural Operator: Learning Fast and Grid-independent PDE Solvers

2022-07-23 · Björn Lütjens, Catherine H. Crawford, Campbell D Watson, Christopher Hill 외

Numerical simulations in climate, chemistry, or astrophysics are computationally too expensive for uncertainty quantification or parameter-exploration at high-resolution. Reduced-order or surrogate models are multiple or…

Operator learningUncertainty Quantification

Model-Free Design of Control Systems over Wireless Fading Channels

2020-09-03 · Vinicius Lima, Mark Eisen, Konstantinos Gatsis, Alejandro Ribeiro

Wireless control systems replace traditional wired communication with wireless networks to exchange information between actuators, plants and sensors in a control system. The noise in wireless channels renders ideal cont…

Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs

2022-06-06 · Dongsheng Ding, Kaiqing Zhang, Jiali Duan, Tamer Başar 외

We study sequential decision making problems aimed at maximizing the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted…

Decision MakingSequential Decision Making

Nonlinear Optimal Control of Electron Dynamics within Hartree-Fock Theory

2024-12-04 · Harish S. Bhat, Hardeep Bassi, Christine M. Isborn

Consider the problem of determining the optimal applied electric field to drive a molecule from an initial state to a desired target state. For even moderately sized molecules, solving this problem directly using the exa…

On the Global Optimum Convergence of Momentum-based Policy Gradient

2021-10-19 · Yuhao Ding, Junzi Zhang, Javad Lavaei

Policy gradient (PG) methods are popular and efficient for large-scale reinforcement learning due to their relative stability and incremental nature. In recent years, the empirical success of PG methods has led to the de…