paper-with-me

Papers

Logit-Q Dynamics for Efficient Learning in Stochastic Teams

2023-02-20 · Ahmed Said Donmez, Onur Unlu, Muhammed O. Sayin

We present a new family of logit-Q dynamics for efficient learning in stochastic games by combining the log-linear learning (also known as logit dynamics) for the repeated play of normal-form games with Q-learning for unknown Markov decision processes within the auxiliary stage-game framework. In this framework, we view stochastic games as agents repeatedly playing some stage game associated with the current state of the underlying game while the agents' Q-functions determine the payoffs of these stage games. We show that the logit-Q dynamics presented reach (near) efficient equilibrium in stochastic teams with unknown dynamics and quantify the approximation error. We also show the rationality of the logit-Q dynamics against agents following pure stationary strategies and the convergence of the dynamics in stochastic games where the stage-payoffs induce potential games, yet only a single agent controls the state transitions beyond stochastic teams. The key idea is to approximate the dynamics with a fictional scenario where the Q-function estimates are stationary over epochs whose lengths grow at a sufficiently slow rate. We then couple the dynamics in the main and fictional scenarios to show that these two scenarios become more and more similar across epochs due to the vanishing step size and growing epoch lengths.

📄 PDF Abstract BibTeX arXiv:2302.09806

Code (0)

등록된 구현이 없습니다.

Tasks

Q-Learning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Endogenous Barriers to Learning

2023-06-29 · Olivier Compte

Building on the idea that lack of experience is a source of errors but that experience should reduce them, we model agents' behavior using a stochastic choice model (logit quantal response), leaving endogenous the accura…

Learning Coordination Policies over Heterogeneous Graphs for Human-Robot Teams via Recurrent Neural Schedule Propagation

2023-01-30 · Batuhan Altundas, Zheyuan Wang, Joshua Bishop, Matthew Gombolay

As human-robot collaboration increases in the workforce, it becomes essential for human-robot teams to coordinate efficiently and intuitively. Traditional approaches for human-robot scheduling either utilize exact method…

Decision MakingGraph AttentionSchedulingSequential Decision Making

Behavioral Foundations of Nested Stochastic Choice and Nested Logit

2021-12-14 · Matthew Kovach, Gerelt Tserenjigmid

We provide the first behavioral characterization of nested logit, a foundational and widely applied discrete choice model, through the introduction of a non-parametric version of nested logit that we call Nested Stochast…

Team Power Dynamics and Team Impact: New Perspectives on Scientific Collaboration using Career Age as a Proxy for Team Power

2021-08-09 · Huimin Xu, Yi Bu, MeiJun Liu, Chenwei Zhang 외

Power dynamics influence every aspect of scientific collaboration. Team power dynamics can be measured by team power level and team power hierarchy. Team power level is conceptualized as the average level of the possessi…

Decision MakingSociology

Efficiency and Stability in a Process of Teams Formation

2021-03-25 · Leonardo Boncinelli, Alessio Muscillo, Paolo Pin

Motivated by data on coauthorships in scientific publications, we analyze a team formation process that generalizes matching models and network formation models, allowing for overlapping teams of heterogeneous size. We a…