paper
-with-
me
Papers
Browse State-of-the-Art
Datasets
Methods
AI Agents
Trends
Digest
🌙
Atari Games
벤치마크
Atari Games on Atari 2600 Double Dunk
86개 결과 ·
⬇ CSV
·
JSON
Score
-18.1
-7.575
2.95
13.48
24
2012-07
2026-09
UCT — 24.0 (2012-07-19)
Best Learner — -13.1 (2012-07-19)
UCT — 24.0 (2012-07-19)
Best Learner — -13.1 (2012-07-19)
Nature DQN — -18.1 (2015-02-25)
Nature DQN — -18.1 (2015-02-25)
Gorila — -11.3 (2015-07-15)
Gorila — -11.3 (2015-07-15)
DQN hs — -6.0 (2015-09-22)
DQN noop — -6.6 (2015-09-22)
DDQN (tuned) hs — -0.3 (2015-09-22)
Prior+Duel hs — -10.7 (2015-09-22)
DQN hs — -6.0 (2015-09-22)
DQN noop — -6.6 (2015-09-22)
DDQN (tuned) hs — -0.3 (2015-09-22)
Prior+Duel hs — -10.7 (2015-09-22)
Prior noop — 18.5 (2015-11-18)
Prior hs — 16.0 (2015-11-18)
Prior noop — 18.5 (2015-11-18)
Prior hs — 16.0 (2015-11-18)
Duel noop — 0.1 (2015-11-20)
Duel hs — -0.8 (2015-11-20)
DDQN (tuned) noop — -5.5 (2015-11-20)
Prior+Duel noop — -12.5 (2015-11-20)
Duel noop — 0.1 (2015-11-20)
Duel hs — -0.8 (2015-11-20)
DDQN (tuned) noop — -5.5 (2015-11-20)
Prior+Duel noop — -12.5 (2015-11-20)
Advantage Learning — -0.15 (2015-12-15)
Persistent AL — -2.51 (2015-12-15)
Advantage Learning — -0.15 (2015-12-15)
Persistent AL — -2.51 (2015-12-15)
A3C FF (1 day) hs — 0.1 (2016-02-04)
A3C LSTM hs — 0.1 (2016-02-04)
A3C FF hs — -0.1 (2016-02-04)
A3C FF (1 day) hs — 0.1 (2016-02-04)
A3C LSTM hs — 0.1 (2016-02-04)
A3C FF hs — -0.1 (2016-02-04)
Bootstrapped DQN — 3.0 (2016-02-15)
Bootstrapped DQN — 3.0 (2016-02-15)
DDQN+Pop-Art noop — -11.5 (2016-02-24)
DDQN+Pop-Art noop — -11.5 (2016-02-24)
ES FF (1 hour) noop — 0.2 (2017-03-10)
ES FF (1 hour) noop — 0.2 (2017-03-10)
Reactor 500M — 23.0 (2017-04-15)
Reactor 500M — 23.0 (2017-04-15)
NoisyNet-Dueling — 1.0 (2017-06-30)
NoisyNet-Dueling — 1.0 (2017-06-30)
C51 noop — 2.5 (2017-07-21)
C51 noop — 2.5 (2017-07-21)
QR-DQN-1 — 21.9 (2017-10-27)
QR-DQN-1 — 21.9 (2017-10-27)
IMPALA (deep) — -0.33 (2018-02-05)
IMPALA (deep) — -0.33 (2018-02-05)
Ape-X — 23.5 (2018-03-02)
Ape-X — 23.5 (2018-03-02)
A2C + SIL — 21.5 (2018-06-14)
IQN — 5.6 (2018-06-14)
CGP — 2.0 (2018-06-14)
A2C + SIL — 21.5 (2018-06-14)
IQN — 5.6 (2018-06-14)
CGP — 2.0 (2018-06-14)
POP3D — -7.89 (2018-07-02)
POP3D — -7.89 (2018-07-02)
R2D2 — 23.7 (2019-05-01)
R2D2 — 23.7 (2019-05-01)
MuZero — 23.94 (2019-11-19)
MuZero — 23.94 (2019-11-19)
Agent57 — 23.93 (2020-03-30)
Agent57 — 23.93 (2020-03-30)
DreamerV2 — 17.0 (2020-10-05)
DreamerV2 — 17.0 (2020-10-05)
MuZero (Res2 Adam) — 23.91 (2021-04-13)
MuZero (Res2 Adam) — 23.91 (2021-04-13)
GDI-H3 — 24.0 (2021-06-11)
GDI-H3 — 24.0 (2021-06-11)
GDI-I3 — 24.0 (2022-06-07)
GDI-H3 — 24.0 (2022-06-07)
GDI-I3 — 24.0 (2022-06-07)
GDI-H3 — 24.0 (2022-06-07)
DNA — -1.3 (2022-06-20)
DNA — -1.3 (2022-06-20)
ASL DDQN — 0.1 (2023-05-07)
ASL DDQN — 0.1 (2023-05-07)
UCT — 24.0 (2012-07-19)
2012-07-19 — UCT: Score 24.0
Rank
Model
Score
Paper
Code
Year
1
UCT
24
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
1
GDI-H3
24
GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
2021
1
GDI-I3
24
Generalized Data Distribution Iteration
2022
1
GDI-H3
24
Generalized Data Distribution Iteration
2022
5
MuZero
23.94
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
werner-duvaud/muzero-general
·
opendilab/LightZero
·
koulanurag/muzero-pytorch
·
+15
2019
6
Agent57
23.93
Agent57: Outperforming the Atari Human Benchmark
michaelnny/deep_rl_zoo
·
pocokhc/agent57
·
yuta0821/agent57_pytorch
·
+2
2020
7
MuZero (Res2 Adam)
23.91
Online and Offline Reinforcement Learning by Planning with a Learned Model
DHDev0/Muzero-unplugged
·
enpasos/muzero
2021
8
R2D2
23.7
Recurrent Experience Replay in Distributed Reinforcement Learning
opendilab/DI-engine
·
michaelnny/deep_rl_zoo
·
garymm/earl
2019
9
Ape-X
23.5
Distributed Prioritized Experience Replay
ray-project/ray
·
vwxyzjn/cleanrl
·
opendilab/DI-engine
·
+12
2018
10
Reactor 500M
23.0
The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning
2017
11
QR-DQN-1
21.9
Distributional Reinforcement Learning with Quantile Regression
DLR-RM/stable-baselines3
·
facebookresearch/Horizon
·
facebookresearch/ReAgent
·
+14
2017
12
A2C + SIL
21.5
Self-Imitation Learning
junhyukoh/self-imitation-learning
·
rwightman/pytorch-opensim-rl
·
SeungeonBaek/continuous-agents-test
·
+1
2018
13
Prior noop
18.5
Prioritized Experience Replay
labmlai/annotated_deep_learning_paper_implementations
·
hill-a/stable-baselines
·
NervanaSystems/coach
·
+74
2015
14
DreamerV2
17
Mastering Atari with Discrete World Models
opendilab/DI-engine
·
danijar/dreamerv2
·
andrejorsula/drl_grasping
·
+6
2020
15
Prior hs
16.0
Prioritized Experience Replay
labmlai/annotated_deep_learning_paper_implementations
·
hill-a/stable-baselines
·
NervanaSystems/coach
·
+74
2015
16
IQN
5.6
Implicit Quantile Networks for Distributional Reinforcement Learning
opendilab/DI-engine
·
chainer/chainerrl
·
Kchu/DeepRL_CK
·
+16
2018
17
Bootstrapped DQN
3
Deep Exploration via Bootstrapped DQN
tensorflow/models
·
tensorflow/models
·
NervanaSystems/coach
·
+3
2016
18
C51 noop
2.5
A Distributional Perspective on Reinforcement Learning
facebookresearch/Horizon
·
facebookresearch/ReAgent
·
opendilab/DI-engine
·
+19
2017
19
CGP
2
Evolving simple programs for playing Atari games
ShuhuaGao/gpFlappyBird
·
JacobLaney/cgp-tetris
2018
20
NoisyNet-Dueling
1
Noisy Networks for Exploration
opendilab/DI-engine
·
Curt-Park/rainbow-is-all-you-need
·
chainer/chainerrl
·
+12
2017
1–20 / 86
다음 →
페이지당
10
20
50
100