paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Double Dunk

86개 결과 · ⬇ CSV · JSON

Score

-18.1 -7.575 2.95 13.48 24 2012-07 2026-09 UCT — 24.0 (2012-07-19) Best Learner — -13.1 (2012-07-19) UCT — 24.0 (2012-07-19) Best Learner — -13.1 (2012-07-19) Nature DQN — -18.1 (2015-02-25) Nature DQN — -18.1 (2015-02-25) Gorila — -11.3 (2015-07-15) Gorila — -11.3 (2015-07-15) DQN hs — -6.0 (2015-09-22) DQN noop — -6.6 (2015-09-22) DDQN (tuned) hs — -0.3 (2015-09-22) Prior+Duel hs — -10.7 (2015-09-22) DQN hs — -6.0 (2015-09-22) DQN noop — -6.6 (2015-09-22) DDQN (tuned) hs — -0.3 (2015-09-22) Prior+Duel hs — -10.7 (2015-09-22) Prior noop — 18.5 (2015-11-18) Prior hs — 16.0 (2015-11-18) Prior noop — 18.5 (2015-11-18) Prior hs — 16.0 (2015-11-18) Duel noop — 0.1 (2015-11-20) Duel hs — -0.8 (2015-11-20) DDQN (tuned) noop — -5.5 (2015-11-20) Prior+Duel noop — -12.5 (2015-11-20) Duel noop — 0.1 (2015-11-20) Duel hs — -0.8 (2015-11-20) DDQN (tuned) noop — -5.5 (2015-11-20) Prior+Duel noop — -12.5 (2015-11-20) Advantage Learning — -0.15 (2015-12-15) Persistent AL — -2.51 (2015-12-15) Advantage Learning — -0.15 (2015-12-15) Persistent AL — -2.51 (2015-12-15) A3C FF (1 day) hs — 0.1 (2016-02-04) A3C LSTM hs — 0.1 (2016-02-04) A3C FF hs — -0.1 (2016-02-04) A3C FF (1 day) hs — 0.1 (2016-02-04) A3C LSTM hs — 0.1 (2016-02-04) A3C FF hs — -0.1 (2016-02-04) Bootstrapped DQN — 3.0 (2016-02-15) Bootstrapped DQN — 3.0 (2016-02-15) DDQN+Pop-Art noop — -11.5 (2016-02-24) DDQN+Pop-Art noop — -11.5 (2016-02-24) ES FF (1 hour) noop — 0.2 (2017-03-10) ES FF (1 hour) noop — 0.2 (2017-03-10) Reactor 500M — 23.0 (2017-04-15) Reactor 500M — 23.0 (2017-04-15) NoisyNet-Dueling — 1.0 (2017-06-30) NoisyNet-Dueling — 1.0 (2017-06-30) C51 noop — 2.5 (2017-07-21) C51 noop — 2.5 (2017-07-21) QR-DQN-1 — 21.9 (2017-10-27) QR-DQN-1 — 21.9 (2017-10-27) IMPALA (deep) — -0.33 (2018-02-05) IMPALA (deep) — -0.33 (2018-02-05) Ape-X — 23.5 (2018-03-02) Ape-X — 23.5 (2018-03-02) A2C + SIL — 21.5 (2018-06-14) IQN — 5.6 (2018-06-14) CGP — 2.0 (2018-06-14) A2C + SIL — 21.5 (2018-06-14) IQN — 5.6 (2018-06-14) CGP — 2.0 (2018-06-14) POP3D — -7.89 (2018-07-02) POP3D — -7.89 (2018-07-02) R2D2 — 23.7 (2019-05-01) R2D2 — 23.7 (2019-05-01) MuZero — 23.94 (2019-11-19) MuZero — 23.94 (2019-11-19) Agent57 — 23.93 (2020-03-30) Agent57 — 23.93 (2020-03-30) DreamerV2 — 17.0 (2020-10-05) DreamerV2 — 17.0 (2020-10-05) MuZero (Res2 Adam) — 23.91 (2021-04-13) MuZero (Res2 Adam) — 23.91 (2021-04-13) GDI-H3 — 24.0 (2021-06-11) GDI-H3 — 24.0 (2021-06-11) GDI-I3 — 24.0 (2022-06-07) GDI-H3 — 24.0 (2022-06-07) GDI-I3 — 24.0 (2022-06-07) GDI-H3 — 24.0 (2022-06-07) DNA — -1.3 (2022-06-20) DNA — -1.3 (2022-06-20) ASL DDQN — 0.1 (2023-05-07) ASL DDQN — 0.1 (2023-05-07) UCT — 24.0 (2012-07-19)
RankModel Score PaperCodeYear
1 UCT 24 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
1 GDI-H3 24 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
1 GDI-I3 24 Generalized Data Distribution Iteration 2022
1 GDI-H3 24 Generalized Data Distribution Iteration 2022
5 MuZero 23.94 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
6 Agent57 23.93 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
7 MuZero (Res2 Adam) 23.91 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
8 R2D2 23.7 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
9 Ape-X 23.5 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
10 Reactor 500M 23.0 The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 2017
11 QR-DQN-1 21.9 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
12 A2C + SIL 21.5 Self-Imitation Learning junhyukoh/self-imitation-learning · rwightman/pytorch-opensim-rl · SeungeonBaek/continuous-agents-test · +1 2018
13 Prior noop 18.5 Prioritized Experience Replay labmlai/annotated_deep_learning_paper_implementations · hill-a/stable-baselines · NervanaSystems/coach · +74 2015
14 DreamerV2 17 Mastering Atari with Discrete World Models opendilab/DI-engine · danijar/dreamerv2 · andrejorsula/drl_grasping · +6 2020
15 Prior hs 16.0 Prioritized Experience Replay labmlai/annotated_deep_learning_paper_implementations · hill-a/stable-baselines · NervanaSystems/coach · +74 2015
16 IQN 5.6 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
17 Bootstrapped DQN 3 Deep Exploration via Bootstrapped DQN tensorflow/models · tensorflow/models · NervanaSystems/coach · +3 2016
18 C51 noop 2.5 A Distributional Perspective on Reinforcement Learning facebookresearch/Horizon · facebookresearch/ReAgent · opendilab/DI-engine · +19 2017
19 CGP 2 Evolving simple programs for playing Atari games ShuhuaGao/gpFlappyBird · JacobLaney/cgp-tetris 2018
20 NoisyNet-Dueling 1 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
1–20 / 86 다음 → 페이지당 10 20 50 100