Conservative and Risk-Aware Offline Multi-Agent Reinforcement Learning
Reinforcement learning (RL) has been widely adopted for controlling and optimizing complex engineering systems such as next-generation wireless networks. An important challenge in adopting RL is the need for direct access to the physical environment. This limitation is particularly severe in multi-agent systems, for which conventional multi-agent reinforcement learning (MARL) requires a large number of coordinated online interactions with the environment during training. When only offline data is available, a direct application of online MARL schemes would generally fail due to the epistemic uncertainty entailed by the lack of exploration during training. In this work, we propose an offline MARL scheme that integrates distributional RL and conservative Q-learning to address the environment's inherent aleatoric uncertainty and the epistemic uncertainty arising from the use of offline data. We explore both independent and joint learning strategies. The proposed MARL scheme, referred to as multi-agent conservative quantile regression, addresses general risk-sensitive design criteria and is applied to the trajectory planning problem in drone networks, showcasing its advantages.
Code (1)
Tasks
Multi-agent Reinforcement LearningQ-Learningquantile regressionreinforcement-learningReinforcement LearningReinforcement Learning (RL)Trajectory PlanningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Conservative Offline Policy Adaptation in Multi-Agent Games
Prior research on policy adaptation in multi-agent games has often relied on online interaction with the target agent in training, which can be expensive and impractical in real-world scenarios. Inspired by recent progre…
Conservative Offline Distributional Reinforcement Learning
Many reinforcement learning (RL) problems in practice are offline, learning purely from observational data. A key challenge is how to ensure the learned policy is safe, which requires quantifying the risk associated with…
D4RLDistributional Reinforcement LearningMuJoCoOffline RL+4Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement Learning
Offline multi-agent reinforcement learning is challenging due to the coupling effect of both distribution shift issue common in offline setting and the high dimension issue common in multi-agent setting, making the actio…
counterfactualMulti-agent Reinforcement LearningOffline RLQ-Learning+2Few is More: Task-Efficient Skill-Discovery for Multi-Task Offline Multi-Agent Reinforcement Learning
As a data-driven approach, offline MARL learns superior policies solely from offline datasets, ideal for domains rich in historical data but with high interaction costs and risks. However, most existing methods are task-…
Learning to ExecuteMulti-agent Reinforcement LearningQ-LearningPlan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification
Conservatism has led to significant progress in offline reinforcement learning (RL) where an agent learns from pre-collected datasets. However, as many real-world scenarios involve interaction among multiple agents, it i…
Continuous ControlMulti-agent Reinforcement LearningOffline RLreinforcement-learning+1