paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Boxing

90개 결과 · ⬇ CSV · JSON

Score

4.8 28.6 52.4 76.2 100 2012-07 2026-09 UCT — 100.0 (2012-07-19) Best Learner — 44.0 (2012-07-19) UCT — 100.0 (2012-07-19) Best Learner — 44.0 (2012-07-19) Nature DQN — 71.8 (2015-02-25) Nature DQN — 71.8 (2015-02-25) Gorila — 74.2 (2015-07-15) Gorila — 74.2 (2015-07-15) DQN noop — 88.0 (2015-09-22) Prior+Duel hs — 79.2 (2015-09-22) DDQN (tuned) hs — 73.5 (2015-09-22) DQN hs — 70.3 (2015-09-22) DQN noop — 88.0 (2015-09-22) Prior+Duel hs — 79.2 (2015-09-22) DDQN (tuned) hs — 73.5 (2015-09-22) DQN hs — 70.3 (2015-09-22) Prior noop — 95.6 (2015-11-18) Prior hs — 72.3 (2015-11-18) Prior noop — 95.6 (2015-11-18) Prior hs — 72.3 (2015-11-18) Duel noop — 99.4 (2015-11-20) Prior+Duel noop — 98.9 (2015-11-20) DDQN (tuned) noop — 91.6 (2015-11-20) Duel hs — 77.3 (2015-11-20) Duel noop — 99.4 (2015-11-20) Prior+Duel noop — 98.9 (2015-11-20) DDQN (tuned) noop — 91.6 (2015-11-20) Duel hs — 77.3 (2015-11-20) Persistent AL — 94.3 (2015-12-15) Advantage Learning — 93.94 (2015-12-15) Persistent AL — 94.3 (2015-12-15) Advantage Learning — 93.94 (2015-12-15) A3C FF hs — 59.8 (2016-02-04) A3C LSTM hs — 37.3 (2016-02-04) A3C FF (1 day) hs — 33.7 (2016-02-04) A3C FF hs — 59.8 (2016-02-04) A3C LSTM hs — 37.3 (2016-02-04) A3C FF (1 day) hs — 33.7 (2016-02-04) Bootstrapped DQN — 93.2 (2016-02-15) Bootstrapped DQN — 93.2 (2016-02-15) DDQN+Pop-Art noop — 99.3 (2016-02-24) DDQN+Pop-Art noop — 99.3 (2016-02-24) ES FF (1 hour) noop — 49.8 (2017-03-10) ES FF (1 hour) noop — 49.8 (2017-03-10) Reactor 500M — 99.4 (2017-04-15) Reactor 500M — 99.4 (2017-04-15) NoisyNet-Dueling — 100.0 (2017-06-30) NoisyNet-Dueling — 100.0 (2017-06-30) C51 noop — 97.8 (2017-07-21) C51 noop — 97.8 (2017-07-21) QR-DQN-1 — 99.9 (2017-10-27) QR-DQN-1 — 99.9 (2017-10-27) DDRL A3C — 98.0 (2018-01-09) DDRL A3C — 98.0 (2018-01-09) IMPALA (deep) — 99.96 (2018-02-05) IMPALA (deep) — 99.96 (2018-02-05) Ape-X — 100.0 (2018-03-02) Ape-X — 100.0 (2018-03-02) IQN — 99.8 (2018-06-14) A2C + SIL — 99.6 (2018-06-14) CGP — 38.4 (2018-06-14) IQN — 99.8 (2018-06-14) A2C + SIL — 99.6 (2018-06-14) CGP — 38.4 (2018-06-14) POP3D — 97.23 (2018-07-02) POP3D — 97.23 (2018-07-02) R2D2 — 98.5 (2019-05-01) R2D2 — 98.5 (2019-05-01) MuZero — 100.0 (2019-11-19) MuZero — 100.0 (2019-11-19) Agent57 — 100.0 (2020-03-30) Agent57 — 100.0 (2020-03-30) CURL — 4.8 (2020-04-08) CURL — 4.8 (2020-04-08) DreamerV2 — 92.0 (2020-10-05) DreamerV2 — 92.0 (2020-10-05) MuZero (Res2 Adam) — 100.0 (2021-04-13) MuZero (Res2 Adam) — 100.0 (2021-04-13) GDI-H3 — 100.0 (2021-06-11) GDI-H3 — 100.0 (2021-06-11) GDI-I3 — 100.0 (2022-06-07) GDI-H3 — 100.0 (2022-06-07) GDI-I3 — 100.0 (2022-06-07) GDI-H3 — 100.0 (2022-06-07) DNA — 99.9 (2022-06-20) DNA — 99.9 (2022-06-20) ASL DDQN — 99.6 (2023-05-07) ASL DDQN — 99.6 (2023-05-07) UCT — 100.0 (2012-07-19)
RankModel Score PaperCodeYear
1 MuZero 100.00 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
1 Ape-X 100 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
1 NoisyNet-Dueling 100 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
1 UCT 100 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
1 Agent57 100 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
1 MuZero (Res2 Adam) 100 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
1 GDI-H3 100 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
1 GDI-I3 100 Generalized Data Distribution Iteration 2022
1 GDI-H3 100 Generalized Data Distribution Iteration 2022
10 IMPALA (deep) 99.96 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
11 QR-DQN-1 99.9 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
11 DNA 99.9 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
13 IQN 99.8 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
14 A2C + SIL 99.6 Self-Imitation Learning junhyukoh/self-imitation-learning · rwightman/pytorch-opensim-rl · SeungeonBaek/continuous-agents-test · +1 2018
14 ASL DDQN 99.6 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
16 Duel noop 99.4 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
16 Reactor 500M 99.4 The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 2017
18 DDQN+Pop-Art noop 99.3 Learning values across many orders of magnitude 2016
19 Prior+Duel noop 98.9 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
20 R2D2 98.5 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
21 DDRL A3C 98 Distributed Deep Reinforcement Learning: Learn how to play Atari games in 21 minutes deepsense-ai/Distributed-BA3C 2018
22 C51 noop 97.8 A Distributional Perspective on Reinforcement Learning facebookresearch/Horizon · facebookresearch/ReAgent · opendilab/DI-engine · +19 2017
23 POP3D 97.23 Policy Optimization With Penalized Point Probability Distance: An Alternative To Proximal Policy Optimization cxxgtxy/POP3D · paperwithcode/pop3d 2018
24 Prior noop 95.6 Prioritized Experience Replay labmlai/annotated_deep_learning_paper_implementations · hill-a/stable-baselines · NervanaSystems/coach · +74 2015
25 Persistent AL 94.3 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
26 Advantage Learning 93.94 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
27 Bootstrapped DQN 93.2 Deep Exploration via Bootstrapped DQN tensorflow/models · tensorflow/models · NervanaSystems/coach · +3 2016
28 DreamerV2 92 Mastering Atari with Discrete World Models opendilab/DI-engine · danijar/dreamerv2 · andrejorsula/drl_grasping · +6 2020
29 DDQN (tuned) noop 91.6 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
30 DQN noop 88.0 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
31 Prior+Duel hs 79.2 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
32 Duel hs 77.3 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
33 Gorila 74.2 Massively Parallel Methods for Deep Reinforcement Learning nandomp/AICollaboratory · londoed/Kortex · londoed/Gorila 2015
34 DDQN (tuned) hs 73.5 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
35 Prior hs 72.3 Prioritized Experience Replay labmlai/annotated_deep_learning_paper_implementations · hill-a/stable-baselines · NervanaSystems/coach · +74 2015
36 Nature DQN 71.8 Human level control through deep reinforcement learning MaximeVandegar/Papers-in-100-Lines-of-Code · gordicaleksa/pytorch-learn-reinforcement-learning · xiuyu0000/new_papers_codes · +5 2015
37 DQN hs 70.3 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
38 A3C FF hs 59.8 Asynchronous Methods for Deep Reinforcement Learning ray-project/ray · DLR-RM/stable-baselines3 · tensorpack/tensorpack · +67 2016
39 ES FF (1 hour) noop 49.8 Evolution Strategies as a Scalable Alternative to Reinforcement Learning ray-project/ray · openai/evolution-strategies-starter · atgambardella/pytorch-es · +20 2017
40 Best Learner 44 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
41 CGP 38.4 Evolving simple programs for playing Atari games ShuhuaGao/gpFlappyBird · JacobLaney/cgp-tetris 2018
42 A3C LSTM hs 37.3 Asynchronous Methods for Deep Reinforcement Learning ray-project/ray · DLR-RM/stable-baselines3 · tensorpack/tensorpack · +67 2016
43 A3C FF (1 day) hs 33.7 Asynchronous Methods for Deep Reinforcement Learning ray-project/ray · DLR-RM/stable-baselines3 · tensorpack/tensorpack · +67 2016
44 SARSA 9.8
45 CURL 4.8 CURL: Contrastive Unsupervised Representations for Reinforcement Learning opendilab/DI-engine · MishaLaskin/curl · aravindsrinivas/curl_rainbow · +4 2020
46 MuZero 100.00 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
46 Ape-X 100 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
46 NoisyNet-Dueling 100 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
46 UCT 100 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
46 Agent57 100 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
1–50 / 90 다음 → 페이지당 10 20 50 100