paper-with-me

홈 › Papers

Co2PO: Coordinated Constrained Policy Optimization for Multi-Agent RL

2026-02-03 · Shrenik Patel, Christine Truong arxiv

Constrained multi-agent reinforcement learning (MARL) faces a fundamental tension between exploration and safety-constrained optimization. Existing leading approaches, such as Lagrangian methods, typically rely on global penalties or centralized critics that react to violations after they occur, often suppressing exploration and leading to over-conservatism. We propose Co2PO, a novel MARL communication-augmented framework that enables coordination-driven safety through selective, risk-aware communication. Co2PO introduces a shared blackboard architecture for broadcasting positional intent and yield signals, governed by a learned hazard predictor that proactively forecasts potential violations over an extended temporal horizon. By integrating these forecasts into a constrained optimization objective, Co2PO allows agents to anticipate and navigate collective hazards without the performance trade-offs inherent in traditional reactive constraints. We evaluate Co2PO across a suite of complex multi-agent safety benchmarks, where it achieves higher returns compared to leading constrained baselines while converging to cost-compliant policies at deployment. Ablation studies further validate the necessity of risk-triggered communication, adaptive gating, and shared memory components.

📄 PDF Abstract BibTeX arXiv:2602.02970

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

Coordinated Proximal Policy Optimization

2021-11-07 · NeurIPS 2021 12 · Zifan Wu, Chao Yu, Deheng Ye, Junge Zhang 외

We present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step s…

StarcraftStarcraft II

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

2025-08-14 · Chak Lam Shek, Guangyao Shi, Pratap Tokekar arxiv

Multi-agent reinforcement learning (MARL) requires coordinated and stable policy updates among interacting agents. Heterogeneous-Agent Trust Region Policy Optimization (HATRPO) enforces per-agent trust region constraints…

Multi-agent Reinforcement Learning

Co-Optimization of Environment and Policies for Decentralized Multi-Agent Navigation

2024-03-21 · Zhan Gao, Guang Yang, Amanda Prorok

This work views the multi-agent system and its surrounding environment as a co-evolving system, where the behavior of one affects the other. The goal is to take both agent actions and environment configurations as decisi…

Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning

2026-06-12 · Pengxin Wang, Lihao Guo, Yi Xie, Bo Liu 외 arxiv

Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives. In this setting, conflicts arise not only across objectives but also…

Multi-agent Reinforcement LearningDecision Making

Coordinated Pandemic Control with Large Language Model Agents as Policymaking Assistants

2026-01-14 · Ziyi Shi, Xusen Guo, Hongliang Lu, Mingxing Peng 외 arxiv

Effective pandemic control requires timely and coordinated policymaking across administrative regions that are intrinsically interdependent. However, human-driven responses are often fragmented and reactive, with policie…