paper-with-me

홈 › Papers

The Power of Exploiter: Provable Multi-Agent RL in Large State Spaces

2021-06-07 · Chi Jin, Qinghua Liu, Tiancheng Yu

Modern reinforcement learning (RL) commonly engages practical problems with large state spaces, where function approximation must be deployed to approximate either the value function or the policy. While recent progresses in RL theory address a rich set of RL problems with general function approximation, such successes are mostly restricted to the single-agent setting. It remains elusive how to extend these results to multi-agent RL, especially due to the new challenges arising from its game-theoretical nature. This paper considers two-player zero-sum Markov Games (MGs). We propose a new algorithm that can provably find the Nash equilibrium policy using a polynomial number of samples, for any MG with low multi-agent Bellman-Eluder dimension -- a new complexity measure adapted from its single-agent version (Jin et al., 2021). A key component of our new algorithm is the exploiter, which facilitates the learning of the main player by deliberately exploiting her weakness. Our theoretical framework is generic, which applies to a wide range of models including but not limited to tabular MGs, MGs with linear or kernel function approximation, and MGs with rich observations.

📄 PDF Abstract BibTeX arXiv:2106.03352

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Minimax Exploiter: A Data Efficient Approach for Competitive Self-Play

2023-11-28 · Daniel Bairamian, Philippe Marcotte, Joshua Romoff, Gabriel Robert 외

Recent advances in Competitive Self-Play (CSP) have achieved, or even surpassed, human level performance in complex game environments such as Dota 2 and StarCraft II using Distributed Multi-Agent Reinforcement Learning (…

Atari GamesDiversityDota 2Multi-agent Reinforcement Learning+2

RoCo: Role-Based LLMs Collaboration for Automatic Heuristic Design

2025-12-03 · Jiawei Xu, Feng-Feng Wei, Wei-Neng Chen arxiv

Automatic Heuristic Design (AHD) has gained traction as a promising solution for solving combinatorial optimization problems (COPs). Large Language Models (LLMs) have emerged and become a promising approach to achieving …

A Robust and Opponent-Aware League Training Method for StarCraft II

2023-09-21 · NeurIPS 2023 11

It is extremely difficult to train a superhuman Artificial Intelligence (AI) for games of similar size to StarCraft II. AlphaStar is the first AI that beat human professionals in the full game of StarCraft II, using a le…

SEREN: Knowing When to Explore and When to Exploit

2022-05-30 · Changmin Yu, David Mguni, Dong Li, Aivar Sootla 외

Efficient reinforcement learning (RL) involves a trade-off between "exploitative" actions that maximise expected reward and "explorative'" ones that sample unvisited states. To encourage exploration, recent approaches pr…

MuJoCoReinforcement Learning (RL)

A Deep Reinforcement Learning Approach for Finding Non-Exploitable Strategies in Two-Player Atari Games

2022-07-18 · Zihan Ding, DiJia Su, Qinghua Liu, Chi Jin

This paper proposes new, end-to-end deep reinforcement learning algorithms for learning two-player zero-sum Markov games. Different from prior efforts on training agents to beat a fixed set of opponents, our objective is…

Atari GamesDeep Reinforcement LearningQ-Learning