paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Skiing

46개 결과 · ⬇ CSV · JSON

Score

-3.002e+04 -2.252e+04 -1.501e+04 -7505 0 2012-07 2026-09 Best Learner — 0.0 (2012-07-19) Full Tree — 0.0 (2012-07-19) Best Learner — 0.0 (2012-07-19) Full Tree — 0.0 (2012-07-19) Advantage Learning — -13264.51 (2015-12-15) Advantage Learning — -13264.51 (2015-12-15) NoisyNet-Dueling — -7550.0 (2017-06-30) NoisyNet-Dueling — -7550.0 (2017-06-30) QR-DQN-1 — -9324.0 (2017-10-27) QR-DQN-1 — -9324.0 (2017-10-27) IMPALA (deep) — -10180.38 (2018-02-05) IMPALA (deep) — -10180.38 (2018-02-05) Ape-X — -10789.9 (2018-03-02) Ape-X — -10789.9 (2018-03-02) IQN — -9289.0 (2018-06-14) CGP — -9011.0 (2018-06-14) IQN — -9289.0 (2018-06-14) CGP — -9011.0 (2018-06-14) R2D2 — -30021.7 (2019-05-01) R2D2 — -30021.7 (2019-05-01) FQF — -9085.3 (2019-11-05) FQF — -9085.3 (2019-11-05) MuZero — -29968.36 (2019-11-19) MuZero — -29968.36 (2019-11-19) Agent57 — -4202.6 (2020-03-30) Agent57 — -4202.6 (2020-03-30) Go-Explore — -3660.0 (2020-04-27) Go-Explore — -3660.0 (2020-04-27) DreamerV2 — -9299.0 (2020-10-05) DreamerV2 — -9299.0 (2020-10-05) Recurrent Rational DQN Average — -23582.0 (2021-02-18) Rational DQN Average — -23487.0 (2021-02-18) Recurrent Rational DQN Average — -23582.0 (2021-02-18) Rational DQN Average — -23487.0 (2021-02-18) MuZero (Res2 Adam) — -30000.0 (2021-04-13) MuZero (Res2 Adam) — -30000.0 (2021-04-13) GDI-I3 — -6774.0 (2021-06-11) GDI-I3 — -6774.0 (2021-06-11) GDI-I3 — -6774.0 (2022-06-07) GDI-H3 — -6025.0 (2022-06-07) GDI-I3 — -6774.0 (2022-06-07) GDI-H3 — -6025.0 (2022-06-07) DNA — -29974.0 (2022-06-20) DNA — -29974.0 (2022-06-20) ASL DDQN — -8295.4 (2023-05-07) ASL DDQN — -8295.4 (2023-05-07) Best Learner — 0.0 (2012-07-19)
RankModel Score Extra Training Data PaperCodeYear
1 Best Learner 0 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
1 Full Tree 0 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
3 IQN -9289 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
4 MuZero -29968.36 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
5 R2D2 -30021.7 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
6 IMPALA (deep) -10180.38 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
7 FQF -9085.3 Fully Parameterized Quantile Function for Distributional Reinforcement Learning opendilab/DI-engine · ku2482/fqf-iqn-qrdqn.pytorch · ku2482/rljax · +3 2019
8 NoisyNet-Dueling -7550 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
9 CGP -9011 Evolving simple programs for playing Atari games ShuhuaGao/gpFlappyBird · JacobLaney/cgp-tetris 2018
10 QR-DQN-1 -9324 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
11 Ape-X -10789.9 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
12 Advantage Learning -13264.51 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
13 Go-Explore -3660 First return, then explore uber-research/go-explore · qgallouedec/lge 2020
14 Agent57 -4202.6 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
15 DreamerV2 -9299 Mastering Atari with Discrete World Models opendilab/DI-engine · danijar/dreamerv2 · andrejorsula/drl_grasping · +6 2020
16 Recurrent Rational DQN Average -23582 Adaptive Rational Activations to Boost Deep Reinforcement Learning ml-research/rational_activations · ml-research/rational_rl · k4ntz/activation-functions · +1 2021
17 Rational DQN Average -23487 Adaptive Rational Activations to Boost Deep Reinforcement Learning ml-research/rational_activations · ml-research/rational_rl · k4ntz/activation-functions · +1 2021
18 MuZero (Res2 Adam) -30000 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
19 GDI-I3 -6774 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
19 GDI-I3 -6774 Generalized Data Distribution Iteration 2022
21 GDI-H3 -6025 Generalized Data Distribution Iteration 2022
22 DNA -29974 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
23 ASL DDQN -8295.4 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
24 Best Learner 0 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
24 Full Tree 0 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
26 IQN -9289 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
27 MuZero -29968.36 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
28 R2D2 -30021.7 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
29 IMPALA (deep) -10180.38 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
30 FQF -9085.3 Fully Parameterized Quantile Function for Distributional Reinforcement Learning opendilab/DI-engine · ku2482/fqf-iqn-qrdqn.pytorch · ku2482/rljax · +3 2019
31 NoisyNet-Dueling -7550 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
32 CGP -9011 Evolving simple programs for playing Atari games ShuhuaGao/gpFlappyBird · JacobLaney/cgp-tetris 2018
33 QR-DQN-1 -9324 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
34 Ape-X -10789.9 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
35 Advantage Learning -13264.51 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
36 Go-Explore -3660 First return, then explore uber-research/go-explore · qgallouedec/lge 2020
37 Agent57 -4202.6 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
38 DreamerV2 -9299 Mastering Atari with Discrete World Models opendilab/DI-engine · danijar/dreamerv2 · andrejorsula/drl_grasping · +6 2020
39 Recurrent Rational DQN Average -23582 Adaptive Rational Activations to Boost Deep Reinforcement Learning ml-research/rational_activations · ml-research/rational_rl · k4ntz/activation-functions · +1 2021
40 Rational DQN Average -23487 Adaptive Rational Activations to Boost Deep Reinforcement Learning ml-research/rational_activations · ml-research/rational_rl · k4ntz/activation-functions · +1 2021
41 MuZero (Res2 Adam) -30000 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
42 GDI-I3 -6774 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
42 GDI-I3 -6774 Generalized Data Distribution Iteration 2022
44 GDI-H3 -6025 Generalized Data Distribution Iteration 2022
45 DNA -29974 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
46 ASL DDQN -8295.4 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
1–46 / 46 페이지당 10 20 50 100