paper-with-me

홈 › Papers

Can Deep Reinforcement Learning Solve Erdos-Selfridge-Spencer Games?

2017-11-07 · ICML 2018 7 · Maithra Raghu, Alex Irpan, Jacob Andreas, Robert Kleinberg, Quoc V. Le, Jon Kleinberg

Deep reinforcement learning has achieved many recent successes, but our understanding of its strengths and limitations is hampered by the lack of rich environments in which we can fully characterize optimal behavior, and correspondingly diagnose individual actions against such a characterization. Here we consider a family of combinatorial games, arising from work of Erdos, Selfridge, and Spencer, and we propose their use as environments for evaluating and comparing different approaches to reinforcement learning. These games have a number of appealing features: they are challenging for current learning approaches, but they form (i) a low-dimensional, simply parametrized environment where (ii) there is a linear closed form solution for optimal behavior from any state, and (iii) the difficulty of the game can be tuned by changing environment parameters in an interpretable way. We use these Erdos-Selfridge-Spencer games not only to compare different algorithms, but test for generalization, make comparisons to supervised learning, analyse multiagent play, and even develop a self play algorithm. Code can be found at: https://github.com/rubai5/ESS_Game

📄 PDF Abstract BibTeX arXiv:1711.02301

Code (1)

rubai5/ESS_Game 공식 구현 tf

Tasks

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

SPENCER: Self-Adaptive Model Distillation for Efficient Code Retrieval

2025-08-01 · Wenchao Gu, Zongyi Lyu, Yanlin Wang, Hongyu Zhang 외 arxiv

Code retrieval aims to provide users with desired code snippets based on users' natural language queries. With the development of deep learning technologies, adopting pre-trained models for this task has become mainstrea…

Natural Language Queries

A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games

2022-06-12 · Samuel Sokota, Ryan D'Orazio, J. Zico Kolter, Nicolas Loizou 외

This work studies an algorithm, which we call magnetic mirror descent, that is inspired by mirror descent and the non-Euclidean proximal gradient algorithm. Our contribution is demonstrating the virtues of magnetic mirro…

Deep Reinforcement LearningMuJoCo Gamesreinforcement-learningReinforcement Learning+1

Optimizing Warfarin Dosing using Deep Reinforcement Learning

2022-02-07 · Sadjad Anzabi Zadeh, W. Nick Street, Barrett W. Thomas

Warfarin is a widely used anticoagulant, and has a narrow therapeutic range. Dosing of warfarin should be individualized, since slight overdosing or underdosing can have catastrophic or even fatal consequences. Despite m…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Point Process Modeling of Drug Overdoses with Heterogeneous and Missing Data

2020-10-12 · Xueying Liu, Jeremy Carter, Brad Ray, George Mohler

Opioid overdose rates have increased in the United States over the past decade and reflect a major public health crisis. Modeling and prediction of drug and opioid hotspots, where a high percentage of events fall in a sm…

ClusteringPoint Processes

Quantifying Local Randomness in Human DNA and RNA Sequences Using Erdos Motifs

2018-09-29

In 1932, Paul Erdos asked whether a random walk constructed from a binary sequence can achieve the lowest possible deviation (lowest discrepancy), for the sequence itself and for all its subsequences formed by homogeneou…