paper-with-me

Papers

Schedule Based Temporal Difference Algorithms

2021-11-23 · Rohan Deb, Meet Gandhi, Shalabh Bhatnagar

Learning the value function of a given policy from data samples is an important problem in Reinforcement Learning. TD($\lambda$) is a popular class of algorithms to solve this problem. However, the weights assigned to different $n$-step returns in TD($\lambda$), controlled by the parameter $\lambda$, decrease exponentially with increasing $n$. In this paper, we present a $\lambda$-schedule procedure that generalizes the TD($\lambda$) algorithm to the case when the parameter $\lambda$ could vary with time-step. This allows flexibility in weight assignment, i.e., the user can specify the weights assigned to different $n$-step returns by choosing a sequence $\{\lambda_t\}_{t \geq 1}$. Based on this procedure, we propose an on-policy algorithm - TD($\lambda$)-schedule, and two off-policy algorithms - GTD($\lambda$)-schedule and TDC($\lambda$)-schedule, respectively. We provide proofs of almost sure convergence for all three algorithms under a general Markov noise framework.

📄 PDF Abstract BibTeX arXiv:2111.11768

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Workload Schedulers -- Genesis, Algorithms and Differences

2025-11-13 · Leszek Sliwko, Vladimir Getov arxiv

This paper presents a novel approach to categorization of modern workload schedulers. We provide descriptions of three classes of schedulers: Operating Systems Process Schedulers, Cluster Systems Jobs Schedulers and Big …

ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling

2025-11-06 · Caelan Garrett, Fabio Ramos arxiv

Bimanual and humanoid robots are appealing because of their human-like ability to leverage multiple arms to efficiently complete tasks. However, controlling multiple arms at once is computationally challenging due to the…

Motion Planning

The Practimum-Optimum Algorithm for Manufacturing Scheduling: A Paradigm Shift Leading to Breakthroughs in Scale and Performance

2024-08-19 · Moshe BenBassat

The Practimum-Optimum (P-O) algorithm represents a paradigm shift in developing automatic optimization products for complex real-life business problems such as large-scale manufacturing scheduling. It leverages deep busi…

Scheduling

Dynamic Trip-Vehicle Dispatch with Scheduled and On-Demand Requests

2019-07-20 · Taoan Huang, Bohui Fang, Xiaohui Bei, Fei Fang

Transportation service providers that dispatch drivers and vehicles to riders start to support both on-demand ride requests posted in real time and rides scheduled in advance, leading to new challenges which, to the best…

Balancing hydrogen delivery in national energy systems: impact of the temporal flexibility of hydrogen delivery on export prices

2025-04-15 · Hazem Abdel-Khalek, Eddy Jalbout, Caspar Schauß, Benjamin Pfluger

Hydrogen is expected to play a key role in the energy transition. Analyses exploring the price of hydrogen usually calculate average or marginal production costs regardless of the time of delivery. A key factor that affe…