paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Space Invaders

110개 결과 · ⬇ CSV · JSON

Score

160.8 3.872e+04 7.727e+04 1.158e+05 1.544e+05 2012-07 2026-09 UCT — 2718.0 (2012-07-19) Best Learner — 250.1 (2012-07-19) UCT — 2718.0 (2012-07-19) Best Learner — 250.1 (2012-07-19) DQN Best — 1075.0 (2013-12-19) DQN Best — 1075.0 (2013-12-19) Nature DQN — 1976.0 (2015-02-25) Nature DQN — 1976.0 (2015-02-25) Gorila — 1183.3 (2015-07-15) Gorila — 1183.3 (2015-07-15) Prior+Duel hs — 8978.0 (2015-09-22) DDQN (tuned) hs — 2628.7 (2015-09-22) DQN noop — 1692.3 (2015-09-22) DQN hs — 1293.8 (2015-09-22) Prior+Duel hs — 8978.0 (2015-09-22) DDQN (tuned) hs — 2628.7 (2015-09-22) DQN noop — 1692.3 (2015-09-22) DQN hs — 1293.8 (2015-09-22) Prior hs — 3912.1 (2015-11-18) Prior noop — 2865.8 (2015-11-18) Prior hs — 3912.1 (2015-11-18) Prior noop — 2865.8 (2015-11-18) Prior+Duel noop — 15311.5 (2015-11-20) Duel noop — 6427.3 (2015-11-20) Duel hs — 5993.1 (2015-11-20) DDQN (tuned) noop — 2525.5 (2015-11-20) Prior+Duel noop — 15311.5 (2015-11-20) Duel noop — 6427.3 (2015-11-20) Duel hs — 5993.1 (2015-11-20) DDQN (tuned) noop — 2525.5 (2015-11-20) DARQN soft — 650.0 (2015-12-05) DARQN soft — 650.0 (2015-12-05) Advantage Learning — 3460.79 (2015-12-15) Persistent AL — 3277.59 (2015-12-15) Advantage Learning — 3460.79 (2015-12-15) Persistent AL — 3277.59 (2015-12-15) A3C LSTM hs — 23846.0 (2016-02-04) A3C FF hs — 15730.5 (2016-02-04) A3C FF (1 day) hs — 2214.7 (2016-02-04) A3C LSTM hs — 23846.0 (2016-02-04) A3C FF hs — 15730.5 (2016-02-04) A3C FF (1 day) hs — 2214.7 (2016-02-04) Bootstrapped DQN — 2893.0 (2016-02-15) Bootstrapped DQN — 2893.0 (2016-02-15) DDQN+Pop-Art noop — 2589.7 (2016-02-24) DDQN+Pop-Art noop — 2589.7 (2016-02-24) ES FF (1 hour) noop — 678.5 (2017-03-10) ES FF (1 hour) noop — 678.5 (2017-03-10) NoisyNet-Dueling — 5909.0 (2017-06-30) NoisyNet-Dueling — 5909.0 (2017-06-30) C51 noop — 5747.0 (2017-07-21) C51 noop — 5747.0 (2017-07-21) MAC — 1173.1 (2017-09-01) MAC — 1173.1 (2017-09-01) Rainbow — 12629.0 (2017-10-06) Rainbow — 12629.0 (2017-10-06) QR-DQN-1 — 20972.0 (2017-10-27) QR-DQN-1 — 20972.0 (2017-10-27) DDRL A3C — 650.0 (2018-01-09) DDRL A3C — 650.0 (2018-01-09) IMPALA (deep) — 43595.78 (2018-02-05) IMPALA (deep) — 43595.78 (2018-02-05) Ape-X — 54681.0 (2018-03-02) Ape-X — 54681.0 (2018-03-02) IDVQ + DRSC + XNES — 830.0 (2018-06-04) IDVQ + DRSC + XNES — 830.0 (2018-06-04) IQN — 28888.0 (2018-06-14) A2C + SIL — 2951.7 (2018-06-14) CGP — 1001.0 (2018-06-14) IQN — 28888.0 (2018-06-14) A2C + SIL — 2951.7 (2018-06-14) CGP — 1001.0 (2018-06-14) POP3D — 1216.15 (2018-07-02) POP3D — 1216.15 (2018-07-02) R2D2 — 43223.4 (2019-05-01) R2D2 — 43223.4 (2019-05-01) SAC — 160.8 (2019-10-16) SAC — 160.8 (2019-10-16) FQF — 46498.3 (2019-11-05) FQF — 46498.3 (2019-11-05) MuZero — 74335.3 (2019-11-19) MuZero — 74335.3 (2019-11-19) Agent57 — 48680.86 (2020-03-30) Agent57 — 48680.86 (2020-03-30) MFEC — 1990.0 (2020-08-21) MFEC — 1990.0 (2020-08-21) DreamerV2 — 2474.0 (2020-10-05) DreamerV2 — 2474.0 (2020-10-05) Recurrent Rational DQN Average — 1395.0 (2021-02-18) Rational DQN Average — 650.0 (2021-02-18) Recurrent Rational DQN Average — 1395.0 (2021-02-18) Rational DQN Average — 650.0 (2021-02-18) MuZero (Res2 Adam) — 3645.63 (2021-04-13) MuZero (Res2 Adam) — 3645.63 (2021-04-13) GDI-I3 — 140460.0 (2021-06-11) GDI-I3 — 140460.0 (2021-06-11) GDI-H3(200M frames) — 154380.0 (2022-06-07) GDI-H3 — 154380.0 (2022-06-07) GDI-I3 — 140460.0 (2022-06-07) GDI-H3(200M frames) — 154380.0 (2022-06-07) GDI-H3 — 154380.0 (2022-06-07) GDI-I3 — 140460.0 (2022-06-07) DNA — 2731.0 (2022-06-20) DNA — 2731.0 (2022-06-20) ASL DDQN — 21602.0 (2023-05-07) ASL DDQN — 21602.0 (2023-05-07) UCT — 2718.0 (2012-07-19) Prior+Duel hs — 8978.0 (2015-09-22) Prior+Duel noop — 15311.5 (2015-11-20) A3C LSTM hs — 23846.0 (2016-02-04) IMPALA (deep) — 43595.78 (2018-02-05) Ape-X — 54681.0 (2018-03-02) MuZero — 74335.3 (2019-11-19) GDI-I3 — 140460.0 (2021-06-11) GDI-H3(200M frames) — 154380.0 (2022-06-07)
RankModel ScoreBest ScoreReturn PaperCodeYear
49 Rational DQN Average 650 Adaptive Rational Activations to Boost Deep Reinforcement Learning ml-research/rational_activations · ml-research/rational_rl · k4ntz/activation-functions · +1 2021
52 SARSA 267.9
53 Best Learner 250.1 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
54 SAC 160.8 Soft Actor-Critic for Discrete Action Settings p-christ/Deep-Reinforcement-Learning-Algorithms-with-PyTorch · ku2482/sac-discrete.pytorch · keraJLi/rejax · +10 2019
55 IQ-Learn 507 IQ-Learn: Inverse soft-Q Learning for Imitation Div99/IQ-Learn · robfiras/ls-iq · google-deepmind/csil · +2 2021
56 GDI-H3(200M frames) 154380 Generalized Data Distribution Iteration 2022
56 GDI-H3 154380 Generalized Data Distribution Iteration 2022
58 GDI-I3 140460 Generalized Data Distribution Iteration 2022
58 GDI-I3 140460 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
60 MuZero 74335.30 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
61 Ape-X 54681 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
62 Agent57 48680.86 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
63 FQF 46498.3 Fully Parameterized Quantile Function for Distributional Reinforcement Learning opendilab/DI-engine · ku2482/fqf-iqn-qrdqn.pytorch · ku2482/rljax · +3 2019
64 IMPALA (deep) 43595.78 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
65 R2D2 43223.4 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
66 IQN 28888 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
67 A3C LSTM hs 23846.0 Asynchronous Methods for Deep Reinforcement Learning ray-project/ray · DLR-RM/stable-baselines3 · tensorpack/tensorpack · +67 2016
68 ASL DDQN 21602 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
69 QR-DQN-1 20972 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
70 A3C FF hs 15730.5 Asynchronous Methods for Deep Reinforcement Learning ray-project/ray · DLR-RM/stable-baselines3 · tensorpack/tensorpack · +67 2016
71 Prior+Duel noop 15311.5 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
72 Rainbow 12629.0 Rainbow: Combining Improvements in Deep Reinforcement Learning thu-ml/tianshou · facebookresearch/ReAgent · facebookresearch/Horizon · +31 2017
73 Prior+Duel hs 8978.0 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
74 Duel noop 6427.3 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
75 Duel hs 5993.1 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
76 NoisyNet-Dueling 5909 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
77 C51 noop 5747.0 A Distributional Perspective on Reinforcement Learning facebookresearch/Horizon · facebookresearch/ReAgent · opendilab/DI-engine · +19 2017
78 Prior hs 3912.1 Prioritized Experience Replay labmlai/annotated_deep_learning_paper_implementations · hill-a/stable-baselines · NervanaSystems/coach · +74 2015
79 MuZero (Res2 Adam) 3645.63 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
80 Advantage Learning 3460.79 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
81 Persistent AL 3277.59 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
82 A2C + SIL 2951.7 Self-Imitation Learning junhyukoh/self-imitation-learning · rwightman/pytorch-opensim-rl · SeungeonBaek/continuous-agents-test · +1 2018
83 Bootstrapped DQN 2893 Deep Exploration via Bootstrapped DQN tensorflow/models · tensorflow/models · NervanaSystems/coach · +3 2016
84 Prior noop 2865.8 Prioritized Experience Replay labmlai/annotated_deep_learning_paper_implementations · hill-a/stable-baselines · NervanaSystems/coach · +74 2015
85 DNA 2731 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
86 UCT 2718 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
87 DDQN (tuned) hs 2628.7 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
88 DDQN+Pop-Art noop 2589.7 Learning values across many orders of magnitude 2016
89 DDQN (tuned) noop 2525.5 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
90 DreamerV2 2474 Mastering Atari with Discrete World Models opendilab/DI-engine · danijar/dreamerv2 · andrejorsula/drl_grasping · +6 2020
91 A3C FF (1 day) hs 2214.7 Asynchronous Methods for Deep Reinforcement Learning ray-project/ray · DLR-RM/stable-baselines3 · tensorpack/tensorpack · +67 2016
92 MFEC 19902490 Model-Free Episodic Control with State Aggregation 2020
93 Nature DQN 1976.0 Human level control through deep reinforcement learning MaximeVandegar/Papers-in-100-Lines-of-Code · gordicaleksa/pytorch-learn-reinforcement-learning · xiuyu0000/new_papers_codes · +5 2015
94 DQN noop 1692.3 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
95 Recurrent Rational DQN Average 1395 Adaptive Rational Activations to Boost Deep Reinforcement Learning ml-research/rational_activations · ml-research/rational_rl · k4ntz/activation-functions · +1 2021
96 DQN hs 1293.8 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
97 POP3D 1216.15 Policy Optimization With Penalized Point Probability Distance: An Alternative To Proximal Policy Optimization cxxgtxy/POP3D · paperwithcode/pop3d 2018
98 Gorila 1183.3 Massively Parallel Methods for Deep Reinforcement Learning nandomp/AICollaboratory · londoed/Kortex · londoed/Gorila 2015
99 MAC 1173.1 Mean Actor Critic kavosh8/MAC · camall3n/atari-MAC 2017
100 DQN Best 1075 Playing Atari with Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · ray-project/ray · DLR-RM/stable-baselines3 · +109 2013
← 이전 51–100 / 110 다음 → 페이지당 10 20 50 100