paper-with-me

홈 › Papers

On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

2017-04-03 · Bolin Gao, Lacra Pavel

In this paper, we utilize results from convex analysis and monotone operator theory to derive additional properties of the softmax function that have not yet been covered in the existing literature. In particular, we show that the softmax function is the monotone gradient map of the log-sum-exp function. By exploiting this connection, we show that the inverse temperature parameter determines the Lipschitz and co-coercivity properties of the softmax function. We then demonstrate the usefulness of these properties through an application in game-theoretic reinforcement learning.

📄 PDF Abstract BibTeX arXiv:1704.00805

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms

2025-10-27 · Pravin Nair arxiv

The softmax function is a basic operator in machine learning and optimization, used in classification, attention mechanisms, reinforcement learning, game theory, and problems involving log-sum-exp terms. Existing robustn…

Reinforcement Learning

A Convergent Variant of the Boltzmann Softmax Operator in Reinforcement Learning

2018-09-27 · Ling Pan, Qingpeng Cai, Qi Meng, Wei Chen 외

The Boltzmann softmax operator can trade-off well between exploration and exploitation according to current estimation in an exponential weighting scheme, which is a promising way to address the exploration-exploitation …

Atari GamesQ-Learningreinforcement-learningReinforcement Learning+1

Neural Replicator Dynamics

2019-06-01 · Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Remi Munos 외

Policy gradient and actor-critic algorithms form the basis of many commonly used training techniques in deep reinforcement learning. Using these algorithms in multiagent environments poses problems such as nonstationarit…

counterfactualDeep Reinforcement LearningPolicy Gradient MethodsReinforcement Learning

Deep Learning Meets Mechanism Design: Key Results and Some Novel Applications

2024-01-11 · V. Udaya Sankar, Vishisht Srihari Rao, Y. Narahari

Mechanism design is essentially reverse engineering of games and involves inducing a game among strategic agents in a way that the induced game satisfies a set of desired properties in an equilibrium of the game. Desirab…

energy managementFairness

Convergence and Price of Anarchy Guarantees of the Softmax Policy Gradient in Markov Potential Games

2022-06-15 · Dingyang Chen, Qi Zhang, Thinh T. Doan

We study the performance of policy gradient methods for the subclass of Markov games known as Markov potential games (MPGs), which extends the notion of normal-form potential games to the stateful setting and includes th…

Policy Gradient Methods