paper-with-me

홈 › Papers

Quality-Aware Exploration Budget Allocation for Cooperative Multi-Agent Reinforcement Learning

2026-05-03 · Dahyun Oh, Minhyuk Yoon, H. Jin Kim arxiv

Cooperative multi-agent reinforcement learning (MARL) requires agents to discover joint strategies in a combinatorially large state-action space, yet effective coordination configurations are exceedingly rare. Intrinsic motivation, which augments task rewards with novelty bonuses, is a popular approach for driving exploration, but its effectiveness hinges on the exploration intensity $β$, where too large a value overwhelms the task signal and causes coordination collapse, while too small a value prevents discovery of rare strategies. We address two complementary challenges: adapting $β$ globally over training, and allocating the exploration budget across agents whose intrinsic reward signals vary in reliability. Our framework combines a return-conditioned sigmoid schedule (RCB) for global intensity control with a per-agent Reward Signal Quality (RSQ) metric that concentrates the exploration budget on agents with reliable signals. The core insight is that agents receiving noisy intrinsic rewards should explore less aggressively, and this allocation can be determined automatically from signal-to-noise statistics. Successor Distance (SD), a quasimetric intrinsic reward, naturally produces distinguishable per-agent signal quality, completing the framework with convergence and ordering preservation guarantees. On seven cooperative benchmarks (MPE, SMAX, MABrax), our method achieves top-tier returns across all environments.

📄 PDF Abstract BibTeX arXiv:2605.01865

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

CoKV: Optimizing KV Cache Allocation via Cooperative Game

2025-02-21 · Qiheng Sun, Hongwei Zhang, Haocheng Xia, Jiayao Zhang 외

Large language models (LLMs) have achieved remarkable success on various aspects of human life. However, one of the major challenges in deploying these models is the substantial memory consumption required to store key-v…

Expressive mechanisms for equitable rent division on a budget

2019-02-08 · Rodrigo A. Velez

We study the incentive properties of envy-free mechanisms for the allocation of rooms and payments of rent among financially constrained roommates. Each agent reports her values for rooms, her housing earmark (soft budge…

Budget Allocation for Unknown Value Functions in a Lipschitz Space

2025-10-12 · MohammadHossein Bateni, Hossein Esfandiari, Samira HosseinGhorban, Alireza Mirrokni 외 arxiv

Building learning models frequently requires evaluating numerous intermediate models. Examples include models considered during feature selection, model structure search, and parameter tunings. The evaluation of an inter…

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

2026-07-29 · Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen 외 arxiv

Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress,…

Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation

2025-09-30 · Ziniu Li, Congliang Chen, Tianyun Yang, Tian Ding 외 arxiv

Large Language Models (LLMs) can self-improve through reinforcement learning, where they generate trajectories to explore and discover better solutions. However, this exploration process is computationally expensive, oft…

Reinforcement LearningMathematical Reasoning