paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Defender

42개 결과 · ⬇ CSV · JSON

Score

3.064e+04 2.712e+05 5.118e+05 7.524e+05 9.93e+05 2015-11 2026-09 Duel noop — 42214.0 (2015-11-20) Prior+Duel noop — 41324.5 (2015-11-20) Prior+Duel hs — 34415.0 (2015-11-20) Duel noop — 42214.0 (2015-11-20) Prior+Duel noop — 41324.5 (2015-11-20) Prior+Duel hs — 34415.0 (2015-11-20) Persistent AL — 32038.93 (2015-12-15) Advantage Learning — 30643.59 (2015-12-15) Persistent AL — 32038.93 (2015-12-15) Advantage Learning — 30643.59 (2015-12-15) Reactor 500M — 223025.0 (2017-04-15) Reactor 500M — 223025.0 (2017-04-15) NoisyNet-Dueling — 42253.0 (2017-06-30) NoisyNet-Dueling — 42253.0 (2017-06-30) QR-DQN-1 — 47887.0 (2017-10-27) QR-DQN-1 — 47887.0 (2017-10-27) IMPALA (deep) — 185203.0 (2018-02-05) IMPALA (deep) — 185203.0 (2018-02-05) Ape-X — 411943.5 (2018-03-02) Ape-X — 411943.5 (2018-03-02) CGP — 993010.0 (2018-06-14) IQN — 53537.0 (2018-06-14) CGP — 993010.0 (2018-06-14) IQN — 53537.0 (2018-06-14) R2D2 — 665792.0 (2019-05-01) R2D2 — 665792.0 (2019-05-01) MuZero — 839642.95 (2019-11-19) MuZero — 839642.95 (2019-11-19) Agent57 — 677642.78 (2020-03-30) Agent57 — 677642.78 (2020-03-30) MuZero (Res2 Adam) — 557200.75 (2021-04-13) MuZero (Res2 Adam) — 557200.75 (2021-04-13) GDI-I3 — 893110.0 (2021-06-11) GDI-I3 — 893110.0 (2021-06-11) GDI-H3 — 970540.0 (2022-06-07) GDI-I3 — 893110.0 (2022-06-07) GDI-H3 — 970540.0 (2022-06-07) GDI-I3 — 893110.0 (2022-06-07) DNA — 152768.0 (2022-06-20) DNA — 152768.0 (2022-06-20) ASL DDQN — 37026.5 (2023-05-07) ASL DDQN — 37026.5 (2023-05-07) Duel noop — 42214.0 (2015-11-20) Reactor 500M — 223025.0 (2017-04-15) Ape-X — 411943.5 (2018-03-02) CGP — 993010.0 (2018-06-14)
RankModel Score PaperCodeYear
1 CGP 993010 Evolving simple programs for playing Atari games ShuhuaGao/gpFlappyBird · JacobLaney/cgp-tetris 2018
2 GDI-H3 970540 Generalized Data Distribution Iteration 2022
3 GDI-I3 893110 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
3 GDI-I3 893110 Generalized Data Distribution Iteration 2022
5 MuZero 839642.95 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
6 Agent57 677642.78 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
7 R2D2 665792.0 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
8 MuZero (Res2 Adam) 557200.75 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
9 Ape-X 411943.5 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
10 Reactor 500M 223025.0 The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 2017
11 IMPALA (deep) 185203.00 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
12 DNA 152768 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
13 IQN 53537 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
14 QR-DQN-1 47887 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
15 NoisyNet-Dueling 42253 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
16 Duel noop 42214.0 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
17 Prior+Duel noop 41324.5 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
18 ASL DDQN 37026.5 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
19 Prior+Duel hs 34415.0 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
20 Persistent AL 32038.93 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
21 Advantage Learning 30643.59 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
22 CGP 993010 Evolving simple programs for playing Atari games ShuhuaGao/gpFlappyBird · JacobLaney/cgp-tetris 2018
23 GDI-H3 970540 Generalized Data Distribution Iteration 2022
24 GDI-I3 893110 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
24 GDI-I3 893110 Generalized Data Distribution Iteration 2022
26 MuZero 839642.95 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
27 Agent57 677642.78 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
28 R2D2 665792.0 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
29 MuZero (Res2 Adam) 557200.75 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
30 Ape-X 411943.5 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
31 Reactor 500M 223025.0 The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 2017
32 IMPALA (deep) 185203.00 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
33 DNA 152768 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
34 IQN 53537 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
35 QR-DQN-1 47887 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
36 NoisyNet-Dueling 42253 Noisy Networks for Exploration opendilab/DI-engine · Curt-Park/rainbow-is-all-you-need · chainer/chainerrl · +12 2017
37 Duel noop 42214.0 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
38 Prior+Duel noop 41324.5 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
39 ASL DDQN 37026.5 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
40 Prior+Duel hs 34415.0 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
41 Persistent AL 32038.93 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
42 Advantage Learning 30643.59 Increasing the Action Gap: New Operators for Reinforcement Learning janhuenermann/neurojs · chainer/chainerrl 2015
1–42 / 42 페이지당 10 20 50 100