paper-with-me

Atari Games 벤치마크

Atari Games on Atari-57

22개 결과 · ⬇ CSV · JSON

Mean Human Normalized Score

504 2897 5291 7684 1.008e+04 2017-10 2026-09 Rainbow DQN — 873.97 (2017-10-06) Rainbow DQN — 873.97 (2017-10-06) IMPALA, deep — 957.34 (2018-02-05) IMPALA, deep — 957.34 (2018-02-05) R2D2 — 3374.31 (2019-05-01) R2D2 — 3374.31 (2019-05-01) LASER — 1741.36 (2019-09-25) LASER — 1741.36 (2019-09-25) MuZero — 4996.2 (2019-11-19) MuZero — 4996.2 (2019-11-19) M-IQN — 504.0 (2020-07-28) M-IQN — 504.0 (2020-07-28) GDI-H3(200M frames) — 9620.98 (2021-06-11) GDI-H3(200M frames) — 9620.98 (2021-06-11) GDI-H3 — 9620.33 (2022-06-07) GDI-I3 — 7810.1 (2022-06-07) GDI-H3 — 9620.33 (2022-06-07) GDI-I3 — 7810.1 (2022-06-07) LBC — 10077.52 (2023-05-09) LBC — 10077.52 (2023-05-09) Rainbow DQN — 873.97 (2017-10-06) IMPALA, deep — 957.34 (2018-02-05) R2D2 — 3374.31 (2019-05-01) MuZero — 4996.2 (2019-11-19) GDI-H3(200M frames) — 9620.98 (2021-06-11) LBC — 10077.52 (2023-05-09)
RankModel Mean Human Normalized ScoreHuman World Record Breakthrough PaperCodeYear
1 LBC 10077.52%24 Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection 2023
2 GDI-H3(200M frames) 9620.98%22 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
3 GDI-H3 9620.33%22 Generalized Data Distribution Iteration 2022
4 GDI-I3 7810.1%17 Generalized Data Distribution Iteration 2022
5 MuZero 4996.20%19 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
6 R2D2 3374.31%15 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
7 LASER 1741.36%7 Off-Policy Actor-Critic with Shared Experience Replay 2019
8 IMPALA, deep 957.34%3 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
9 Rainbow DQN 873.97%4 Rainbow: Combining Improvements in Deep Reinforcement Learning thu-ml/tianshou · facebookresearch/ReAgent · facebookresearch/Horizon · +31 2017
10 M-IQN 504% Munchausen Reinforcement Learning google-research/google-research · deepmind/acme · opendilab/DI-engine · +3 2020
11 GDI-H3 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
12 LBC 10077.52%24 Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection 2023
13 GDI-H3(200M frames) 9620.98%22 GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 2021
14 GDI-H3 9620.33%22 Generalized Data Distribution Iteration 2022
15 GDI-I3 7810.1%17 Generalized Data Distribution Iteration 2022
16 MuZero 4996.20%19 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
17 R2D2 3374.31%15 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
18 LASER 1741.36%7 Off-Policy Actor-Critic with Shared Experience Replay 2019
19 IMPALA, deep 957.34%3 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
20 Rainbow DQN 873.97%4 Rainbow: Combining Improvements in Deep Reinforcement Learning thu-ml/tianshou · facebookresearch/ReAgent · facebookresearch/Horizon · +31 2017
1–20 / 22 다음 → 페이지당 10 20 50 100