paper
-with-
me
Papers
Browse State-of-the-Art
Datasets
Methods
AI Agents
Trends
Digest
🌙
Atari Games
벤치마크
Atari Games on Atari 2600 Boxing
90개 결과 ·
⬇ CSV
·
JSON
Score
4.8
28.6
52.4
76.2
100
2012-07
2026-09
UCT — 100.0 (2012-07-19)
Best Learner — 44.0 (2012-07-19)
UCT — 100.0 (2012-07-19)
Best Learner — 44.0 (2012-07-19)
Nature DQN — 71.8 (2015-02-25)
Nature DQN — 71.8 (2015-02-25)
Gorila — 74.2 (2015-07-15)
Gorila — 74.2 (2015-07-15)
DQN noop — 88.0 (2015-09-22)
Prior+Duel hs — 79.2 (2015-09-22)
DDQN (tuned) hs — 73.5 (2015-09-22)
DQN hs — 70.3 (2015-09-22)
DQN noop — 88.0 (2015-09-22)
Prior+Duel hs — 79.2 (2015-09-22)
DDQN (tuned) hs — 73.5 (2015-09-22)
DQN hs — 70.3 (2015-09-22)
Prior noop — 95.6 (2015-11-18)
Prior hs — 72.3 (2015-11-18)
Prior noop — 95.6 (2015-11-18)
Prior hs — 72.3 (2015-11-18)
Duel noop — 99.4 (2015-11-20)
Prior+Duel noop — 98.9 (2015-11-20)
DDQN (tuned) noop — 91.6 (2015-11-20)
Duel hs — 77.3 (2015-11-20)
Duel noop — 99.4 (2015-11-20)
Prior+Duel noop — 98.9 (2015-11-20)
DDQN (tuned) noop — 91.6 (2015-11-20)
Duel hs — 77.3 (2015-11-20)
Persistent AL — 94.3 (2015-12-15)
Advantage Learning — 93.94 (2015-12-15)
Persistent AL — 94.3 (2015-12-15)
Advantage Learning — 93.94 (2015-12-15)
A3C FF hs — 59.8 (2016-02-04)
A3C LSTM hs — 37.3 (2016-02-04)
A3C FF (1 day) hs — 33.7 (2016-02-04)
A3C FF hs — 59.8 (2016-02-04)
A3C LSTM hs — 37.3 (2016-02-04)
A3C FF (1 day) hs — 33.7 (2016-02-04)
Bootstrapped DQN — 93.2 (2016-02-15)
Bootstrapped DQN — 93.2 (2016-02-15)
DDQN+Pop-Art noop — 99.3 (2016-02-24)
DDQN+Pop-Art noop — 99.3 (2016-02-24)
ES FF (1 hour) noop — 49.8 (2017-03-10)
ES FF (1 hour) noop — 49.8 (2017-03-10)
Reactor 500M — 99.4 (2017-04-15)
Reactor 500M — 99.4 (2017-04-15)
NoisyNet-Dueling — 100.0 (2017-06-30)
NoisyNet-Dueling — 100.0 (2017-06-30)
C51 noop — 97.8 (2017-07-21)
C51 noop — 97.8 (2017-07-21)
QR-DQN-1 — 99.9 (2017-10-27)
QR-DQN-1 — 99.9 (2017-10-27)
DDRL A3C — 98.0 (2018-01-09)
DDRL A3C — 98.0 (2018-01-09)
IMPALA (deep) — 99.96 (2018-02-05)
IMPALA (deep) — 99.96 (2018-02-05)
Ape-X — 100.0 (2018-03-02)
Ape-X — 100.0 (2018-03-02)
IQN — 99.8 (2018-06-14)
A2C + SIL — 99.6 (2018-06-14)
CGP — 38.4 (2018-06-14)
IQN — 99.8 (2018-06-14)
A2C + SIL — 99.6 (2018-06-14)
CGP — 38.4 (2018-06-14)
POP3D — 97.23 (2018-07-02)
POP3D — 97.23 (2018-07-02)
R2D2 — 98.5 (2019-05-01)
R2D2 — 98.5 (2019-05-01)
MuZero — 100.0 (2019-11-19)
MuZero — 100.0 (2019-11-19)
Agent57 — 100.0 (2020-03-30)
Agent57 — 100.0 (2020-03-30)
CURL — 4.8 (2020-04-08)
CURL — 4.8 (2020-04-08)
DreamerV2 — 92.0 (2020-10-05)
DreamerV2 — 92.0 (2020-10-05)
MuZero (Res2 Adam) — 100.0 (2021-04-13)
MuZero (Res2 Adam) — 100.0 (2021-04-13)
GDI-H3 — 100.0 (2021-06-11)
GDI-H3 — 100.0 (2021-06-11)
GDI-I3 — 100.0 (2022-06-07)
GDI-H3 — 100.0 (2022-06-07)
GDI-I3 — 100.0 (2022-06-07)
GDI-H3 — 100.0 (2022-06-07)
DNA — 99.9 (2022-06-20)
DNA — 99.9 (2022-06-20)
ASL DDQN — 99.6 (2023-05-07)
ASL DDQN — 99.6 (2023-05-07)
UCT — 100.0 (2012-07-19)
2012-07-19 — UCT: Score 100.0
Rank
Model
Score
Paper
Code
Year
1
MuZero
100.00
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
werner-duvaud/muzero-general
·
opendilab/LightZero
·
koulanurag/muzero-pytorch
·
+15
2019
1
Ape-X
100
Distributed Prioritized Experience Replay
ray-project/ray
·
vwxyzjn/cleanrl
·
opendilab/DI-engine
·
+12
2018
1
NoisyNet-Dueling
100
Noisy Networks for Exploration
opendilab/DI-engine
·
Curt-Park/rainbow-is-all-you-need
·
chainer/chainerrl
·
+12
2017
1
UCT
100
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
1
Agent57
100
Agent57: Outperforming the Atari Human Benchmark
michaelnny/deep_rl_zoo
·
pocokhc/agent57
·
yuta0821/agent57_pytorch
·
+2
2020
1
MuZero (Res2 Adam)
100
Online and Offline Reinforcement Learning by Planning with a Learned Model
DHDev0/Muzero-unplugged
·
enpasos/muzero
2021
1
GDI-H3
100
GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
2021
1
GDI-I3
100
Generalized Data Distribution Iteration
2022
1
GDI-H3
100
Generalized Data Distribution Iteration
2022
10
IMPALA (deep)
99.96
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
ray-project/ray
·
opendilab/DI-engine
·
deepmind/haiku
·
+21
2018
11
QR-DQN-1
99.9
Distributional Reinforcement Learning with Quantile Regression
DLR-RM/stable-baselines3
·
facebookresearch/Horizon
·
facebookresearch/ReAgent
·
+14
2017
11
DNA
99.9
DNA: Proximal Policy Optimization with a Dual Network Architecture
maitchison/PPO
2022
13
IQN
99.8
Implicit Quantile Networks for Distributional Reinforcement Learning
opendilab/DI-engine
·
chainer/chainerrl
·
Kchu/DeepRL_CK
·
+16
2018
14
A2C + SIL
99.6
Self-Imitation Learning
junhyukoh/self-imitation-learning
·
rwightman/pytorch-opensim-rl
·
SeungeonBaek/continuous-agents-test
·
+1
2018
14
ASL DDQN
99.6
Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity
xinjinghao/color
2023
16
Duel noop
99.4
Dueling Network Architectures for Deep Reinforcement Learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
facebookresearch/Horizon
·
+70
2015
16
Reactor 500M
99.4
The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning
2017
18
DDQN+Pop-Art noop
99.3
Learning values across many orders of magnitude
2016
19
Prior+Duel noop
98.9
Dueling Network Architectures for Deep Reinforcement Learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
facebookresearch/Horizon
·
+70
2015
20
R2D2
98.5
Recurrent Experience Replay in Distributed Reinforcement Learning
opendilab/DI-engine
·
michaelnny/deep_rl_zoo
·
garymm/earl
2019
21
DDRL A3C
98
Distributed Deep Reinforcement Learning: Learn how to play Atari games in 21 minutes
deepsense-ai/Distributed-BA3C
2018
22
C51 noop
97.8
A Distributional Perspective on Reinforcement Learning
facebookresearch/Horizon
·
facebookresearch/ReAgent
·
opendilab/DI-engine
·
+19
2017
23
POP3D
97.23
Policy Optimization With Penalized Point Probability Distance: An Alternative To Proximal Policy Optimization
cxxgtxy/POP3D
·
paperwithcode/pop3d
2018
24
Prior noop
95.6
Prioritized Experience Replay
labmlai/annotated_deep_learning_paper_implementations
·
hill-a/stable-baselines
·
NervanaSystems/coach
·
+74
2015
25
Persistent AL
94.3
Increasing the Action Gap: New Operators for Reinforcement Learning
janhuenermann/neurojs
·
chainer/chainerrl
2015
26
Advantage Learning
93.94
Increasing the Action Gap: New Operators for Reinforcement Learning
janhuenermann/neurojs
·
chainer/chainerrl
2015
27
Bootstrapped DQN
93.2
Deep Exploration via Bootstrapped DQN
tensorflow/models
·
tensorflow/models
·
NervanaSystems/coach
·
+3
2016
28
DreamerV2
92
Mastering Atari with Discrete World Models
opendilab/DI-engine
·
danijar/dreamerv2
·
andrejorsula/drl_grasping
·
+6
2020
29
DDQN (tuned) noop
91.6
Dueling Network Architectures for Deep Reinforcement Learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
facebookresearch/Horizon
·
+70
2015
30
DQN noop
88.0
Deep Reinforcement Learning with Double Q-learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
hill-a/stable-baselines
·
+94
2015
31
Prior+Duel hs
79.2
Deep Reinforcement Learning with Double Q-learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
hill-a/stable-baselines
·
+94
2015
32
Duel hs
77.3
Dueling Network Architectures for Deep Reinforcement Learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
facebookresearch/Horizon
·
+70
2015
33
Gorila
74.2
Massively Parallel Methods for Deep Reinforcement Learning
nandomp/AICollaboratory
·
londoed/Kortex
·
londoed/Gorila
2015
34
DDQN (tuned) hs
73.5
Deep Reinforcement Learning with Double Q-learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
hill-a/stable-baselines
·
+94
2015
35
Prior hs
72.3
Prioritized Experience Replay
labmlai/annotated_deep_learning_paper_implementations
·
hill-a/stable-baselines
·
NervanaSystems/coach
·
+74
2015
36
Nature DQN
71.8
Human level control through deep reinforcement learning
MaximeVandegar/Papers-in-100-Lines-of-Code
·
gordicaleksa/pytorch-learn-reinforcement-learning
·
xiuyu0000/new_papers_codes
·
+5
2015
37
DQN hs
70.3
Deep Reinforcement Learning with Double Q-learning
labmlai/annotated_deep_learning_paper_implementations
·
tensorpack/tensorpack
·
hill-a/stable-baselines
·
+94
2015
38
A3C FF hs
59.8
Asynchronous Methods for Deep Reinforcement Learning
ray-project/ray
·
DLR-RM/stable-baselines3
·
tensorpack/tensorpack
·
+67
2016
39
ES FF (1 hour) noop
49.8
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
ray-project/ray
·
openai/evolution-strategies-starter
·
atgambardella/pytorch-es
·
+20
2017
40
Best Learner
44
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
41
CGP
38.4
Evolving simple programs for playing Atari games
ShuhuaGao/gpFlappyBird
·
JacobLaney/cgp-tetris
2018
42
A3C LSTM hs
37.3
Asynchronous Methods for Deep Reinforcement Learning
ray-project/ray
·
DLR-RM/stable-baselines3
·
tensorpack/tensorpack
·
+67
2016
43
A3C FF (1 day) hs
33.7
Asynchronous Methods for Deep Reinforcement Learning
ray-project/ray
·
DLR-RM/stable-baselines3
·
tensorpack/tensorpack
·
+67
2016
44
SARSA
9.8
45
CURL
4.8
CURL: Contrastive Unsupervised Representations for Reinforcement Learning
opendilab/DI-engine
·
MishaLaskin/curl
·
aravindsrinivas/curl_rainbow
·
+4
2020
46
MuZero
100.00
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
werner-duvaud/muzero-general
·
opendilab/LightZero
·
koulanurag/muzero-pytorch
·
+15
2019
46
Ape-X
100
Distributed Prioritized Experience Replay
ray-project/ray
·
vwxyzjn/cleanrl
·
opendilab/DI-engine
·
+12
2018
46
NoisyNet-Dueling
100
Noisy Networks for Exploration
opendilab/DI-engine
·
Curt-Park/rainbow-is-all-you-need
·
chainer/chainerrl
·
+12
2017
46
UCT
100
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
46
Agent57
100
Agent57: Outperforming the Atari Human Benchmark
michaelnny/deep_rl_zoo
·
pocokhc/agent57
·
yuta0821/agent57_pytorch
·
+2
2020
1–50 / 90
다음 →
페이지당
10
20
50
100