paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Boxing

90개 결과 · ⬇ CSV · JSON

Score

4.8 28.6 52.4 76.2 100 2012-07 2026-09 UCT — 100.0 (2012-07-19) Best Learner — 44.0 (2012-07-19) UCT — 100.0 (2012-07-19) Best Learner — 44.0 (2012-07-19) Nature DQN — 71.8 (2015-02-25) Nature DQN — 71.8 (2015-02-25) Gorila — 74.2 (2015-07-15) Gorila — 74.2 (2015-07-15) DQN noop — 88.0 (2015-09-22) Prior+Duel hs — 79.2 (2015-09-22) DDQN (tuned) hs — 73.5 (2015-09-22) DQN hs — 70.3 (2015-09-22) DQN noop — 88.0 (2015-09-22) Prior+Duel hs — 79.2 (2015-09-22) DDQN (tuned) hs — 73.5 (2015-09-22) DQN hs — 70.3 (2015-09-22) Prior noop — 95.6 (2015-11-18) Prior hs — 72.3 (2015-11-18) Prior noop — 95.6 (2015-11-18) Prior hs — 72.3 (2015-11-18) Duel noop — 99.4 (2015-11-20) Prior+Duel noop — 98.9 (2015-11-20) DDQN (tuned) noop — 91.6 (2015-11-20) Duel hs — 77.3 (2015-11-20) Duel noop — 99.4 (2015-11-20) Prior+Duel noop — 98.9 (2015-11-20) DDQN (tuned) noop — 91.6 (2015-11-20) Duel hs — 77.3 (2015-11-20) Persistent AL — 94.3 (2015-12-15) Advantage Learning — 93.94 (2015-12-15) Persistent AL — 94.3 (2015-12-15) Advantage Learning — 93.94 (2015-12-15) A3C FF hs — 59.8 (2016-02-04) A3C LSTM hs — 37.3 (2016-02-04) A3C FF (1 day) hs — 33.7 (2016-02-04) A3C FF hs — 59.8 (2016-02-04) A3C LSTM hs — 37.3 (2016-02-04) A3C FF (1 day) hs — 33.7 (2016-02-04) Bootstrapped DQN — 93.2 (2016-02-15) Bootstrapped DQN — 93.2 (2016-02-15) DDQN+Pop-Art noop — 99.3 (2016-02-24) DDQN+Pop-Art noop — 99.3 (2016-02-24) ES FF (1 hour) noop — 49.8 (2017-03-10) ES FF (1 hour) noop — 49.8 (2017-03-10) Reactor 500M — 99.4 (2017-04-15) Reactor 500M — 99.4 (2017-04-15) NoisyNet-Dueling — 100.0 (2017-06-30) NoisyNet-Dueling — 100.0 (2017-06-30) C51 noop — 97.8 (2017-07-21) C51 noop — 97.8 (2017-07-21) QR-DQN-1 — 99.9 (2017-10-27) QR-DQN-1 — 99.9 (2017-10-27) DDRL A3C — 98.0 (2018-01-09) DDRL A3C — 98.0 (2018-01-09) IMPALA (deep) — 99.96 (2018-02-05) IMPALA (deep) — 99.96 (2018-02-05) Ape-X — 100.0 (2018-03-02) Ape-X — 100.0 (2018-03-02) IQN — 99.8 (2018-06-14) A2C + SIL — 99.6 (2018-06-14) CGP — 38.4 (2018-06-14) IQN — 99.8 (2018-06-14) A2C + SIL — 99.6 (2018-06-14) CGP — 38.4 (2018-06-14) POP3D — 97.23 (2018-07-02) POP3D — 97.23 (2018-07-02) R2D2 — 98.5 (2019-05-01) R2D2 — 98.5 (2019-05-01) MuZero — 100.0 (2019-11-19) MuZero — 100.0 (2019-11-19) Agent57 — 100.0 (2020-03-30) Agent57 — 100.0 (2020-03-30) CURL — 4.8 (2020-04-08) CURL — 4.8 (2020-04-08) DreamerV2 — 92.0 (2020-10-05) DreamerV2 — 92.0 (2020-10-05) MuZero (Res2 Adam) — 100.0 (2021-04-13) MuZero (Res2 Adam) — 100.0 (2021-04-13) GDI-H3 — 100.0 (2021-06-11) GDI-H3 — 100.0 (2021-06-11) GDI-I3 — 100.0 (2022-06-07) GDI-H3 — 100.0 (2022-06-07) GDI-I3 — 100.0 (2022-06-07) GDI-H3 — 100.0 (2022-06-07) DNA — 99.9 (2022-06-20) DNA — 99.9 (2022-06-20) ASL DDQN — 99.6 (2023-05-07) ASL DDQN — 99.6 (2023-05-07) UCT — 100.0 (2012-07-19)
RankModel Score PaperCodeYear
1 MuZero 100.00 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
1 Ape-X 100 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
1 NoisyNet-Dueling 100 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
1 UCT 100 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
1 Agent57 100 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
1 MuZero (Res2 Adam) 100 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
1 GDI-H3 100 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
1 GDI-I3 100 Generalized Data Distribution Iteration 2022
1 GDI-H3 100 Generalized Data Distribution Iteration 2022
10 IMPALA (deep) 99.96 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
1–10 / 90 다음 → 페이지당 10 20 50 100