paper-with-me

Atari Games 벤치마크

Atari Games on Atari 2600 Asterix

98개 결과 · ⬇ CSV · JSON

Score

272 2.502e+05 5.001e+05 7.501e+05 1e+06 2012-07 2026-09 UCT — 290700.0 (2012-07-19) Best Learner — 987.3 (2012-07-19) UCT — 290700.0 (2012-07-19) Best Learner — 987.3 (2012-07-19) Nature DQN — 6012.0 (2015-02-25) Nature DQN — 6012.0 (2015-02-25) Gorila — 3324.7 (2015-07-15) Gorila — 3324.7 (2015-07-15) Prior+Duel hs — 364200.0 (2015-09-22) DDQN (tuned) hs — 16837.0 (2015-09-22) DQN noop — 4359.0 (2015-09-22) DQN hs — 3170.5 (2015-09-22) Prior+Duel hs — 364200.0 (2015-09-22) DDQN (tuned) hs — 16837.0 (2015-09-22) DQN noop — 4359.0 (2015-09-22) DQN hs — 3170.5 (2015-09-22) Prior noop — 31527.0 (2015-11-18) Prior hs — 22484.5 (2015-11-18) Prior noop — 31527.0 (2015-11-18) Prior hs — 22484.5 (2015-11-18) Prior+Duel noop — 375080.0 (2015-11-20) Prior+Duel hs — 364200.0 (2015-11-20) Duel noop — 28188.0 (2015-11-20) DDQN (tuned) noop — 17356.5 (2015-11-20) Duel hs — 15840.0 (2015-11-20) Prior+Duel noop — 375080.0 (2015-11-20) Prior+Duel hs — 364200.0 (2015-11-20) Duel noop — 28188.0 (2015-11-20) DDQN (tuned) noop — 17356.5 (2015-11-20) Duel hs — 15840.0 (2015-11-20) Persistent AL — 19564.9 (2015-12-15) Advantage Learning — 12852.08 (2015-12-15) Persistent AL — 19564.9 (2015-12-15) Advantage Learning — 12852.08 (2015-12-15) A3C FF hs — 22140.5 (2016-02-04) A3C LSTM hs — 17244.5 (2016-02-04) A3C FF (1 day) hs — 6723.0 (2016-02-04) A3C FF hs — 22140.5 (2016-02-04) A3C LSTM hs — 17244.5 (2016-02-04) A3C FF (1 day) hs — 6723.0 (2016-02-04) Bootstrapped DQN — 19713.2 (2016-02-15) Bootstrapped DQN — 19713.2 (2016-02-15) DDQN+Pop-Art noop — 18919.5 (2016-02-24) DDQN+Pop-Art noop — 18919.5 (2016-02-24) ES FF (1 hour) noop — 1440.0 (2017-03-10) ES FF (1 hour) noop — 1440.0 (2017-03-10) Reactor 500M — 205914.0 (2017-04-15) Reactor 500M — 205914.0 (2017-04-15) NoisyNet-Dueling — 28350.0 (2017-06-30) NoisyNet-Dueling — 28350.0 (2017-06-30) C51 noop — 406211.0 (2017-07-21) C51 noop — 406211.0 (2017-07-21) QR-DQN-1 — 261025.0 (2017-10-27) QR-DQN-1 — 261025.0 (2017-10-27) IMPALA (deep) — 300732.0 (2018-02-05) IMPALA (deep) — 300732.0 (2018-02-05) Ape-X — 313305.0 (2018-03-02) Ape-X — 313305.0 (2018-03-02) IQN — 342016.0 (2018-06-14) A2C + SIL — 17984.2 (2018-06-14) CGP — 1880.0 (2018-06-14) IQN — 342016.0 (2018-06-14) A2C + SIL — 17984.2 (2018-06-14) CGP — 1880.0 (2018-06-14) POP3D — 4310.67 (2018-07-02) POP3D — 4310.67 (2018-07-02) R2D2 — 999153.3 (2019-05-01) R2D2 — 999153.3 (2019-05-01) SAC — 272.0 (2019-10-16) SAC — 272.0 (2019-10-16) FQF — 578388.5 (2019-11-05) FQF — 578388.5 (2019-11-05) MuZero — 998425.0 (2019-11-19) MuZero — 998425.0 (2019-11-19) Agent57 — 991384.42 (2020-03-30) Agent57 — 991384.42 (2020-03-30) CURL — 524.3 (2020-04-08) CURL — 524.3 (2020-04-08) DreamerV2 — 72311.0 (2020-10-05) DreamerV2 — 72311.0 (2020-10-05) Rational DQN Average — 18109.0 (2021-02-18) Recurrent Rational DQN Average — 12621.0 (2021-02-18) Rational DQN Average — 18109.0 (2021-02-18) Recurrent Rational DQN Average — 12621.0 (2021-02-18) MuZero (Res2 Adam) — 862406.65 (2021-04-13) MuZero (Res2 Adam) — 862406.65 (2021-04-13) GDI-H3 — 999999.0 (2022-06-07) GDI-I3 — 759910.0 (2022-06-07) GDI-H3 — 999999.0 (2022-06-07) GDI-I3 — 759910.0 (2022-06-07) DNA — 323965.0 (2022-06-20) DNA — 323965.0 (2022-06-20) ASL DDQN — 567640.0 (2023-05-07) ASL DDQN — 567640.0 (2023-05-07) UCT — 290700.0 (2012-07-19) Prior+Duel hs — 364200.0 (2015-09-22) Prior+Duel noop — 375080.0 (2015-11-20) C51 noop — 406211.0 (2017-07-21) R2D2 — 999153.3 (2019-05-01) GDI-H3 — 999999.0 (2022-06-07)
RankModel Score PaperCodeYear
1 GDI-H3 999999 Generalized Data Distribution Iteration 2022
2 R2D2 999153.3 Recurrent Experience Replay in Distributed Reinforcement Learning opendilab/DI-engine · michaelnny/deep_rl_zoo · garymm/earl 2019
3 MuZero 998425.00 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model werner-duvaud/muzero-general · opendilab/LightZero · koulanurag/muzero-pytorch · +15 2019
4 Agent57 991384.42 Agent57: Outperforming the Atari Human Benchmark michaelnny/deep_rl_zoo · pocokhc/agent57 · yuta0821/agent57_pytorch · +2 2020
5 MuZero (Res2 Adam) 862406.65 Online and Offline Reinforcement Learning by Planning with a Learned Model DHDev0/Muzero-unplugged · enpasos/muzero 2021
6 GDI-I3 759910 Generalized Data Distribution Iteration 2022
7 FQF 578388.5 Fully Parameterized Quantile Function for Distributional Reinforcement Learning opendilab/DI-engine · ku2482/fqf-iqn-qrdqn.pytorch · ku2482/rljax · +3 2019
8 ASL DDQN 567640 Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity xinjinghao/color 2023
9 C51 noop 406211 A Distributional Perspective on Reinforcement Learning facebookresearch/Horizon · facebookresearch/ReAgent · opendilab/DI-engine · +19 2017
10 Prior+Duel noop 375080.0 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
11 Prior+Duel hs 364200.0 Deep Reinforcement Learning with Double Q-learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · hill-a/stable-baselines · +94 2015
11 Prior+Duel hs 364200.0 Dueling Network Architectures for Deep Reinforcement Learning labmlai/annotated_deep_learning_paper_implementations · tensorpack/tensorpack · facebookresearch/Horizon · +70 2015
13 IQN 342016 Implicit Quantile Networks for Distributional Reinforcement Learning opendilab/DI-engine · chainer/chainerrl · Kchu/DeepRL_CK · +16 2018
14 DNA 323965 DNA: Proximal Policy Optimization with a Dual Network Architecture maitchison/PPO 2022
15 Ape-X 313305 Distributed Prioritized Experience Replay ray-project/ray · vwxyzjn/cleanrl · opendilab/DI-engine · +12 2018
16 IMPALA (deep) 300732.00 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures ray-project/ray · opendilab/DI-engine · deepmind/haiku · +21 2018
17 UCT 290700 The Arcade Learning Environment: An Evaluation Platform for General Agents mgbellemare/Arcade-Learning-Environment · kenjyoung/MinAtar · nandomp/AICollaboratory 2012
18 QR-DQN-1 261025 Distributional Reinforcement Learning with Quantile Regression DLR-RM/stable-baselines3 · facebookresearch/Horizon · facebookresearch/ReAgent · +14 2017
19 Reactor 500M 205914.0 The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 2017
20 DreamerV2 72311 Mastering Atari with Discrete World Models opendilab/DI-engine · danijar/dreamerv2 · andrejorsula/drl_grasping · +6 2020
1–20 / 98 다음 → 페이지당 10 20 50 100