Structured Reinforcement Learning for Delay-Optimal Data Transmission in Dense mmWave Networks
We study the data packet transmission problem (mmDPT) in dense cell-free millimeter wave (mmWave) networks, i.e., users sending data packet requests to access points (APs) via uplinks and APs transmitting requested data packets to users via downlinks. Our objective is to minimize the average delay in the system due to APs' limited service capacity and unreliable wireless channels between APs and users. This problem can be formulated as a restless multi-armed bandits problem with fairness constraint (RMAB-F). Since finding the optimal policy for RMAB-F is intractable, existing learning algorithms are computationally expensive and not suitable for practical dynamic dense mmWave networks. In this paper, we propose a structured reinforcement learning (RL) solution for mmDPT by exploiting the inherent structure encoded in RMAB-F. To achieve this, we first design a low-complexity and provably asymptotically optimal index policy for RMAB-F. Then, we leverage this structure information to develop a structured RL algorithm called mmDPT-TS, which provably achieves an \tilde{O}(\sqrt{T}) Bayesian regret. More importantly, mmDPT-TS is computation-efficient and thus amenable to practical implementation, as it fully exploits the structure of index policy for making decisions. Extensive emulation based on data collected in realistic mmWave networks demonstrate significant gains of mmDPT-TS over existing approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
FairnessMulti-Armed BanditsReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Deep Reinforcement Learning-based Cell DTX/DRX Configuration for Network Energy Saving
3GPP Release 18 cell discontinuous transmission and reception (cell DTX/DRX) is an important new network energy saving feature for 5G. As a time-domain technique, it periodically aggregates the user data transmissions in…
Reinforcement LearningStructured Optimal Transmission Control in Network-coded Two-way Relay Channels
This paper considers a transmission control problem in network-coded two-way relay channels (NC-TWRC), where the relay buffers random symbol arrivals from two users, and the channels are assumed to be fading. The problem…
Vocal Bursts Valence PredictionBalanced Performance Between Energy-Delay and Bit Error Rate in UAV Relay Networks
This paper presents a new strategy for simultaneously reducing energy consumption, transmission delays, and bit error rate in Unmanned Aerial Vehicle UAV networks. A UAV is fitted with a wireless Bidirectional Relay BR t…
Delay-aware Resource Allocation in Fog-assisted IoT Networks Through Reinforcement Learning
Fog nodes in the vicinity of IoT devices are promising to provision low latency services by offloading tasks from IoT devices to them. Mobile IoT is composed by mobile IoT devices such as vehicles, wearable devices and s…
reinforcement-learningReinforcement Learning (RL)Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
We study reinforcement learning with delayed state observation, where the agent observes the current state after some random number of time steps. We propose an algorithm that combines the augmentation method and the upp…
Reinforcement Learning