paper-with-me

홈 › Papers

Game Solving with Online Fine-Tuning

2023-11-13 · NeurIPS 2023 11 · Ti-Rong Wu, Hung Guei, Ting Han Wei, Chung-Chin Shih, Jui-Te Chin, I-Chen Wu

Game solving is a similar, yet more difficult task than mastering a game. Solving a game typically means to find the game-theoretic value (outcome given optimal play), and optionally a full strategy to follow in order to achieve that outcome. The AlphaZero algorithm has demonstrated super-human level play, and its powerful policy and value predictions have also served as heuristics in game solving. However, to solve a game and obtain a full strategy, a winning response must be found for all possible moves by the losing player. This includes very poor lines of play from the losing side, for which the AlphaZero self-play process will not encounter. AlphaZero-based heuristics can be highly inaccurate when evaluating these out-of-distribution positions, which occur throughout the entire search. To address this issue, this paper investigates applying online fine-tuning while searching and proposes two methods to learn tailor-designed heuristics for game solving. Our experiments show that using online fine-tuning can solve a series of challenging 7x7 Killall-Go problems, using only 23.54% of computation time compared to the baseline without online fine-tuning. Results suggest that the savings scale with problem size. Our method can further be extended to any tree search algorithm for problem solving. Our code is available at https://rlg.iis.sinica.edu.tw/papers/neurips2023-online-fine-tuning-solver.

📄 PDF Abstract BibTeX arXiv:2311.07178

Code (1)

rlglab/online-fine-tuning-solver 공식 구현

Tasks

Board Games

Methods 이 논문이 사용한 방법론

AlphaZero AlphaZero is a reinforcement learning agent for playing board games such as Go, chess, and shogi.

Similar Papers 제목 키워드 기반

Reasoning, Memorization, and Fine-Tuning Language Models for Non-Cooperative Games

2024-10-18 · Yunhao Yang, Leonard Berthellemy, Ufuk Topcu

We develop a method that integrates the tree of thoughts and multi-agent framework to enhance the capability of pre-trained language models in solving complex, unfamiliar games. The method decomposes game-solving into fo…

Language ModelingLanguage ModellingMemorization

Safe Subgame Resolving for Extensive Form Correlated Equilibrium

2022-12-29 · Chun Kai Ling, Fei Fang

Correlated Equilibrium is a solution concept that is more general than Nash Equilibrium (NE) and can lead to outcomes with better social welfare. However, its natural extension to the sequential setting, the \textit{Exte…

Form

Scalable Online Planning via Reinforcement Learning Fine-Tuning

2021-09-30 · NeurIPS 2021 12 · Arnaud Fickinger, Hengyuan Hu, Brandon Amos, Stuart Russell 외

Lookahead search has been a critical component of recent AI successes, such as in the games of chess, go, and poker. However, the search methods used in these games, and in many other settings, are tabular. Tabular searc…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Beyond Playtesting: A Generative Multi-Agent Simulation System for Massively Multiplayer Online Games

2025-12-02 · Ran Zhang, Kun Ouyang, Tiancheng Ma, Yida Yang 외 arxiv

Optimizing numerical systems and mechanism design is crucial for enhancing player experience in Massively Multiplayer Online (MMO) games. Traditional optimization approaches rely on large-scale online experiments or para…

Reinforcement Learning

Fine-Tuning Pre-trained Language Models to Detect In-Game Trash Talks

2024-03-19 · Daniel Fesalbon, Arvin De La Cruz, Marvin Mallari, Nelson Rodelas

Common problems in playing online mobile and computer games were related to toxic behavior and abusive communication among players. Based on different reports and studies, the study also discusses the impact of online ha…

Dota 2