paper-with-me

홈 › Papers

The Price of Paranoia: Robust Risk-Sensitive Cooperation in Non-Stationary Multi-Agent Reinforcement Learning

2026-04-17 · Deep Kumar Ganguly, Chandradithya S Jonnalagadda, Pratham Chintamani, Adithya Ananth arxiv

Cooperative equilibria are fragile. When agents learn alongside each other rather than in a fixed environment, the process of learning destabilizes the cooperation they are trying to sustain: every gradient step an agent takes shifts the distribution of actions its partner will play, turning a cooperative partner into a source of stochastic noise precisely where the cooperation decision is most sensitive. We study how this co-learning noise propagates through the structure of coordination games, and find that the cooperative equilibrium, even when strongly Pareto-dominant, is exponentially unstable under standard risk-neutral learning, collapsing irreversibly once partner noise crosses the game's critical cooperation threshold. The natural response to apply distributional robustness to hedge against partner uncertainty makes things strictly worse: risk-averse return objectives penalize the high-variance cooperative action relative to defection, widening the instability region rather than shrinking it, a paradox that reveals a fundamental mismatch between the domains where robustness is applied and instability originates. We resolve this by showing that robustness should target the policy gradient update variance induced by partner uncertainty, not the return distribution. This distinction yields an algorithm whose gradient updates are modulated by an online measure of partner unpredictability, provably expanding the cooperation basin in symmetric coordination games. To unify stability, sample complexity, and welfare consequences of this approach, we introduce the Price of Paranoia as the structural dual of the Price of Anarchy. Together with a novel Cooperation Window, it precisely characterizes how much welfare learning algorithms can recover under partner noise, pinning down the optimal degree of robustness as a closed-form balance between equilibrium stability and sample efficiency.

📄 PDF Abstract BibTeX arXiv:2604.15695

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

Non-stationary Risk-sensitive Reinforcement Learning: Near-optimal Dynamic Regret, Adaptive Detection, and Separation Design

2022-11-19 · Yuhao Ding, Ming Jin, Javad Lavaei

We study risk-sensitive reinforcement learning (RL) based on an entropic risk measure in episodic non-stationary Markov decision processes (MDPs). Both the reward functions and the state transition kernels are unknown an…

Reinforcement Learning (RL)

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

2026-05-08 · Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo 외 arxiv

Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social dilemmas. Across 7 LLMs and 4 games over 500 rounds, expanding accessi…

Stationarity analysis of the stock market data and its application to mechanical trading

2021-12-23 · Kazuki Kanehira, Norikazu Todoroki

This study proposes a scheme for stationarity analysis of stock price fluctuations based on KM$_2$O-Langevin theory. Using this scheme, we classify the time-series data of stock price fluctuations into three periods: sta…

Stock Price PredictionTime SeriesTime Series Analysis

Cost-Sensitive Portfolio Selection via Deep Reinforcement Learning

2020-03-06 · Yifan Zhang, Peilin Zhao, Qingyao Wu, Bin Li 외

Portfolio Selection is an important real-world financial task and has attracted extensive attention in artificial intelligence communities. This task, however, has two main difficulties: (i) the non-stationary price seri…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Can maker-taker fees prevent algorithmic cooperation in market making?

2022-11-01 · Bingyan Han

In a semi-realistic market simulator, independent reinforcement learning algorithms may facilitate market makers to maintain wide spreads even without communication. This unexpected outcome challenges the current antitru…

reinforcement-learningReinforcement Learning (RL)