paper-with-me

홈 › Papers

A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability

2025-06-04 · Wenhan Xu, Jiashuo Jiang, Lei Deng, Danny Hin-Kwok Tsang

With the proliferation of Internet of Things (IoT) devices, the demand for addressing complex optimization challenges has intensified. The Lyapunov Drift-Plus-Penalty algorithm is a widely adopted approach for ensuring queue stability, and some research has preliminarily explored its integration with reinforcement learning (RL). In this paper, we investigate the adaptation of the Lyapunov Drift-Plus-Penalty algorithm for RL applications, deriving an effective method for combining Lyapunov Drift-Plus-Penalty with RL under a set of common and reasonable conditions through rigorous theoretical analysis. Unlike existing approaches that directly merge the two frameworks, our proposed algorithm, termed Lyapunov drift-plus-penalty method tailored for reinforcement learning with queue stability (LDPTRLQ) algorithm, offers theoretical superiority by effectively balancing the greedy optimization of Lyapunov Drift-Plus-Penalty with the long-term perspective of RL. Simulation results for multiple problems demonstrate that LDPTRLQ outperforms the baseline methods using the Lyapunov drift-plus-penalty method and RL, corroborating the validity of our theoretical derivations. The results also demonstrate that our proposed algorithm outperforms other benchmarks in terms of compatibility and stability.

📄 PDF Abstract BibTeX arXiv:2506.04291

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

A Reinforcement Learning Formulation of the Lyapunov Optimization: Application to Edge Computing Systems with Queue Stability

2020-12-14 · Sohee Bae, Seungyul Han, Youngchul Sung

In this paper, a deep reinforcement learning (DRL)-based approach to the Lyapunov optimization is considered to minimize the time-average penalty while maintaining queue stability. A proper construction of state and acti…

Deep Reinforcement LearningEdge-computingreinforcement-learningReinforcement Learning+1

AoI-Delay Tradeoff in Mobile Edge Caching: A Mixed-Order Drift-Plus-Penalty Algorithm

2023-04-18 · Ran Li, Chuan Huang, Xiaoqi Qin, Lei Yang

Mobile edge caching (MEC) is a promising technique to improve the quality of service (QoS) for mobile users (MU) by bringing data to the network edge. However, optimizing the crucial QoS aspects of message freshness and …

Decision MakingSchedulingSequential Decision Making

Goal-oriented Estimation of Multiple Markov Sources in Resource-constrained Systems

2023-11-13 · Jiping Luo, Nikolaos Pappas

This paper investigates goal-oriented communication for remote estimation of multiple Markov sources in resource-constrained networks. An agent decides the updating times of the sources and transmits the packet to a remo…

Deep Reinforcement Learning

Minimizing the AoI in Resource-Constrained Multi-Source Relaying Systems: Dynamic and Learning-based Scheduling

2022-03-10 · Abolfazl Zakeri, Mohammad Moltafet, Markus Leinonen, Marian Codreanu

We consider a multi-source relaying system where independent sources randomly generate status update packets which are sent to the destination with the aid of a relay through unreliable links. We develop transmission sch…

Deep Reinforcement LearningSchedulingStochastic Optimization

Dynamic Joint Communications and Sensing Precoding Design: A Lyapunov Approach

2025-03-18 · Abolfazl Zakeri, Nhan Thanh Nguyen, Ahmed Alkhateeb, Markku Juntti

This letter proposes a dynamic joint communications and sensing (JCAS) framework to adaptively design dedicated sensing and communications precoders. We first formulate a stochastic control problem to maximize the long-t…