paper
-with-
me
Papers
Browse State-of-the-Art
Datasets
Methods
AI Agents
Trends
Digest
🌙
Atari Games
벤치마크
Atari Games on Atari 2600 Boxing
90개 결과 ·
⬇ CSV
·
JSON
Score
4.8
28.6
52.4
76.2
100
2012-07
2026-09
UCT — 100.0 (2012-07-19)
Best Learner — 44.0 (2012-07-19)
UCT — 100.0 (2012-07-19)
Best Learner — 44.0 (2012-07-19)
Nature DQN — 71.8 (2015-02-25)
Nature DQN — 71.8 (2015-02-25)
Gorila — 74.2 (2015-07-15)
Gorila — 74.2 (2015-07-15)
DQN noop — 88.0 (2015-09-22)
Prior+Duel hs — 79.2 (2015-09-22)
DDQN (tuned) hs — 73.5 (2015-09-22)
DQN hs — 70.3 (2015-09-22)
DQN noop — 88.0 (2015-09-22)
Prior+Duel hs — 79.2 (2015-09-22)
DDQN (tuned) hs — 73.5 (2015-09-22)
DQN hs — 70.3 (2015-09-22)
Prior noop — 95.6 (2015-11-18)
Prior hs — 72.3 (2015-11-18)
Prior noop — 95.6 (2015-11-18)
Prior hs — 72.3 (2015-11-18)
Duel noop — 99.4 (2015-11-20)
Prior+Duel noop — 98.9 (2015-11-20)
DDQN (tuned) noop — 91.6 (2015-11-20)
Duel hs — 77.3 (2015-11-20)
Duel noop — 99.4 (2015-11-20)
Prior+Duel noop — 98.9 (2015-11-20)
DDQN (tuned) noop — 91.6 (2015-11-20)
Duel hs — 77.3 (2015-11-20)
Persistent AL — 94.3 (2015-12-15)
Advantage Learning — 93.94 (2015-12-15)
Persistent AL — 94.3 (2015-12-15)
Advantage Learning — 93.94 (2015-12-15)
A3C FF hs — 59.8 (2016-02-04)
A3C LSTM hs — 37.3 (2016-02-04)
A3C FF (1 day) hs — 33.7 (2016-02-04)
A3C FF hs — 59.8 (2016-02-04)
A3C LSTM hs — 37.3 (2016-02-04)
A3C FF (1 day) hs — 33.7 (2016-02-04)
Bootstrapped DQN — 93.2 (2016-02-15)
Bootstrapped DQN — 93.2 (2016-02-15)
DDQN+Pop-Art noop — 99.3 (2016-02-24)
DDQN+Pop-Art noop — 99.3 (2016-02-24)
ES FF (1 hour) noop — 49.8 (2017-03-10)
ES FF (1 hour) noop — 49.8 (2017-03-10)
Reactor 500M — 99.4 (2017-04-15)
Reactor 500M — 99.4 (2017-04-15)
NoisyNet-Dueling — 100.0 (2017-06-30)
NoisyNet-Dueling — 100.0 (2017-06-30)
C51 noop — 97.8 (2017-07-21)
C51 noop — 97.8 (2017-07-21)
QR-DQN-1 — 99.9 (2017-10-27)
QR-DQN-1 — 99.9 (2017-10-27)
DDRL A3C — 98.0 (2018-01-09)
DDRL A3C — 98.0 (2018-01-09)
IMPALA (deep) — 99.96 (2018-02-05)
IMPALA (deep) — 99.96 (2018-02-05)
Ape-X — 100.0 (2018-03-02)
Ape-X — 100.0 (2018-03-02)
IQN — 99.8 (2018-06-14)
A2C + SIL — 99.6 (2018-06-14)
CGP — 38.4 (2018-06-14)
IQN — 99.8 (2018-06-14)
A2C + SIL — 99.6 (2018-06-14)
CGP — 38.4 (2018-06-14)
POP3D — 97.23 (2018-07-02)
POP3D — 97.23 (2018-07-02)
R2D2 — 98.5 (2019-05-01)
R2D2 — 98.5 (2019-05-01)
MuZero — 100.0 (2019-11-19)
MuZero — 100.0 (2019-11-19)
Agent57 — 100.0 (2020-03-30)
Agent57 — 100.0 (2020-03-30)
CURL — 4.8 (2020-04-08)
CURL — 4.8 (2020-04-08)
DreamerV2 — 92.0 (2020-10-05)
DreamerV2 — 92.0 (2020-10-05)
MuZero (Res2 Adam) — 100.0 (2021-04-13)
MuZero (Res2 Adam) — 100.0 (2021-04-13)
GDI-H3 — 100.0 (2021-06-11)
GDI-H3 — 100.0 (2021-06-11)
GDI-I3 — 100.0 (2022-06-07)
GDI-H3 — 100.0 (2022-06-07)
GDI-I3 — 100.0 (2022-06-07)
GDI-H3 — 100.0 (2022-06-07)
DNA — 99.9 (2022-06-20)
DNA — 99.9 (2022-06-20)
ASL DDQN — 99.6 (2023-05-07)
ASL DDQN — 99.6 (2023-05-07)
UCT — 100.0 (2012-07-19)
2012-07-19 — UCT: Score 100.0
Rank
Model
Score
Paper
Code
Year
1
MuZero
100.00
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
werner-duvaud/muzero-general
·
opendilab/LightZero
·
koulanurag/muzero-pytorch
·
+15
2019
1
Ape-X
100
Distributed Prioritized Experience Replay
ray-project/ray
·
vwxyzjn/cleanrl
·
opendilab/DI-engine
·
+12
2018
1
NoisyNet-Dueling
100
Noisy Networks for Exploration
opendilab/DI-engine
·
Curt-Park/rainbow-is-all-you-need
·
chainer/chainerrl
·
+12
2017
1
UCT
100
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
1
Agent57
100
Agent57: Outperforming the Atari Human Benchmark
michaelnny/deep_rl_zoo
·
pocokhc/agent57
·
yuta0821/agent57_pytorch
·
+2
2020
1
MuZero (Res2 Adam)
100
Online and Offline Reinforcement Learning by Planning with a Learned Model
DHDev0/Muzero-unplugged
·
enpasos/muzero
2021
1
GDI-H3
100
GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
2021
1
GDI-I3
100
Generalized Data Distribution Iteration
2022
1
GDI-H3
100
Generalized Data Distribution Iteration
2022
10
IMPALA (deep)
99.96
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
ray-project/ray
·
opendilab/DI-engine
·
deepmind/haiku
·
+21
2018
1–10 / 90
다음 →
페이지당
10
20
50
100