paper-with-me

홈 › Papers

Transformer Based Planning in the Observation Space with Applications to Trick Taking Card Games

2024-04-19 · Douglas Rebstock, Christopher Solinas, Nathan R. Sturtevant, Michael Buro

Traditional search algorithms have issues when applied to games of imperfect information where the number of possible underlying states and trajectories are very large. This challenge is particularly evident in trick-taking card games. While state sampling techniques such as Perfect Information Monte Carlo (PIMC) search has shown success in these contexts, they still have major limitations. We present Generative Observation Monte Carlo Tree Search (GO-MCTS), which utilizes MCTS on observation sequences generated by a game specific model. This method performs the search within the observation space and advances the search using a model that depends solely on the agent's observations. Additionally, we demonstrate that transformers are well-suited as the generative model in this context, and we demonstrate a process for iteratively training the transformer via population-based self-play. The efficacy of GO-MCTS is demonstrated in various games of imperfect information, such as Hearts, Skat, and "The Crew: The Quest for Planet Nine," with promising results.

📄 PDF Abstract BibTeX arXiv:2404.13150

Code (0)

등록된 구현이 없습니다.

Tasks

Card Games

Similar Papers 제목 키워드 기반

Adaptive Online Packing-guided Search for POMDPs

2021-12-01 · NeurIPS 2021 12 · Chenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu 외

The partially observable Markov decision process (POMDP) provides a general framework for modeling an agent's decision process with state uncertainty, and online planning plays a pivotal role in solving it. A belief is a…

Quantum algorithms applied to satellite mission planning for Earth observation

2023-02-14 · Serge Rainjonneau, Igor Tokarev, Sergei Iudin, Saaketh Rayaprolu 외

Earth imaging satellites are a crucial part of our everyday lives that enable global tracking of industrial activities. Use cases span many applications, from weather forecasting to digital maps, carbon footprint trackin…

Earth Observationreinforcement-learningReinforcement Learning (RL)Weather Forecasting

Working Memory Graphs

2019-11-17 · ICML 2020 1 · Ricky Loynd, Roland Fernandez, Asli Celikyilmaz, Adith Swaminathan 외

Transformers have increasingly outperformed gated RNNs in obtaining new state-of-the-art results on supervised tasks involving text sequences. Inspired by this trend, we study the question of how Transformer-based models…

Decision MakingSequential Decision MakingSokoban

Transformer tricks: Precomputing the first layer

2024-02-20 · Nils Graef

This micro-paper describes a trick to speed up inference of transformers with RoPE (such as LLaMA, Mistral, PaLM, and Gemma). For these models, a large portion of the first transformer layer can be precomputed, which res…

RAE-NWM: Navigation World Model in Dense Visual Representation Space

2026-03-10 · Mingkun Zhang, Wangtian Shen, Fan Zhang, Haijian Qin 외 arxiv

Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulating action-conditioned state transitions to predict future observations. …

Visual Navigation