paper-with-me

Papers

Minimax Strikes Back

2020-12-19 · Quentin Cohen-Solal, Tristan Cazenave

Deep Reinforcement Learning reaches a superhuman level of play in many complete information games. The state of the art algorithm for learning with zero knowledge is AlphaZero. We take another approach, Ath\'enan, which uses a different, Minimax-based, search algorithm called Descent, as well as different learning targets and that does not use a policy. We show that for multiple games it is much more efficient than the reimplementation of AlphaZero: Polygames. It is even competitive with Polygames when Polygames uses 100 times more GPU (at least for some games). One of the keys to the superior performance is that the cost of generating state data for training is approximately 296 times lower with Ath\'enan. With the same reasonable ressources, Ath\'enan without reinforcement heuristic is at least 7 times faster than Polygames and much more than 30 times faster with reinforcement heuristic.

📄 PDF Abstract BibTeX arXiv:2012.10700

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningGPUreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Bird Movement Prediction Using Long Short-Term Memory Networks to Prevent Bird Strikes with Low Altitude Aircraft

2023-12-17 · Elaheh Sabziyan Varnousfaderani, Syed A. M. Shihab

The number of collisions between aircraft and birds in the airspace has been increasing at an alarming rate over the past decade due to increasing bird population, air traffic and usage of quieter aircraft. Bird strikes …

Can unions impose costs on employers in education strikes? Evidence from pension disputes in UK universities

2024-01-10 · Nils Braakmann, Barbara Eberth

The impact of strikes in educational institutions, specifically universities, on employers remains understudied. This paper investigates the impact of education strikes in UK universities from 2018 to 2022, primarily due…

Monte Carlo Tree Search with Heuristic Evaluations using Implicit Minimax Backups

2014-06-02 · Marc Lanctot, Mark H. M. Winands, Tom Pepels, Nathan R. Sturtevant

Monte Carlo Tree Search (MCTS) has improved the performance of game engines in domains such as Go, Hex, and general game playing. MCTS has been shown to outperform classic alpha-beta search in games where good heuristic …

On the Minimax Regret in Online Ranking with Top-k Feedback

2023-09-05 · Mingyuan Zhang, Ambuj Tewari

In online ranking, a learning algorithm sequentially ranks a set of items and receives feedback on its ranking in the form of relevance scores. Since obtaining relevance scores typically involves human annotation, it is …

The empire strikes back: Some responses to Bruineberg and colleagues

2021-12-31 · Maxwell J. D. Ramstead

In their target paper, Bruineberg and colleagues provide us with a timely opportunity to discuss the formal constructs and philosophical implications of the free-energy principle. I critically discuss their proposed dist…