paper
-with-
me
Papers
Browse State-of-the-Art
Datasets
Methods
AI Agents
Trends
Digest
🌙
Atari Games
벤치마크
Atari Games on Atari 2600 Skiing
46개 결과 ·
⬇ CSV
·
JSON
Score
-3.002e+04
-2.252e+04
-1.501e+04
-7505
0
2012-07
2026-09
Best Learner — 0.0 (2012-07-19)
Full Tree — 0.0 (2012-07-19)
Best Learner — 0.0 (2012-07-19)
Full Tree — 0.0 (2012-07-19)
Advantage Learning — -13264.51 (2015-12-15)
Advantage Learning — -13264.51 (2015-12-15)
NoisyNet-Dueling — -7550.0 (2017-06-30)
NoisyNet-Dueling — -7550.0 (2017-06-30)
QR-DQN-1 — -9324.0 (2017-10-27)
QR-DQN-1 — -9324.0 (2017-10-27)
IMPALA (deep) — -10180.38 (2018-02-05)
IMPALA (deep) — -10180.38 (2018-02-05)
Ape-X — -10789.9 (2018-03-02)
Ape-X — -10789.9 (2018-03-02)
IQN — -9289.0 (2018-06-14)
CGP — -9011.0 (2018-06-14)
IQN — -9289.0 (2018-06-14)
CGP — -9011.0 (2018-06-14)
R2D2 — -30021.7 (2019-05-01)
R2D2 — -30021.7 (2019-05-01)
FQF — -9085.3 (2019-11-05)
FQF — -9085.3 (2019-11-05)
MuZero — -29968.36 (2019-11-19)
MuZero — -29968.36 (2019-11-19)
Agent57 — -4202.6 (2020-03-30)
Agent57 — -4202.6 (2020-03-30)
Go-Explore — -3660.0 (2020-04-27)
Go-Explore — -3660.0 (2020-04-27)
DreamerV2 — -9299.0 (2020-10-05)
DreamerV2 — -9299.0 (2020-10-05)
Recurrent Rational DQN Average — -23582.0 (2021-02-18)
Rational DQN Average — -23487.0 (2021-02-18)
Recurrent Rational DQN Average — -23582.0 (2021-02-18)
Rational DQN Average — -23487.0 (2021-02-18)
MuZero (Res2 Adam) — -30000.0 (2021-04-13)
MuZero (Res2 Adam) — -30000.0 (2021-04-13)
GDI-I3 — -6774.0 (2021-06-11)
GDI-I3 — -6774.0 (2021-06-11)
GDI-I3 — -6774.0 (2022-06-07)
GDI-H3 — -6025.0 (2022-06-07)
GDI-I3 — -6774.0 (2022-06-07)
GDI-H3 — -6025.0 (2022-06-07)
DNA — -29974.0 (2022-06-20)
DNA — -29974.0 (2022-06-20)
ASL DDQN — -8295.4 (2023-05-07)
ASL DDQN — -8295.4 (2023-05-07)
Best Learner — 0.0 (2012-07-19)
2012-07-19 — Best Learner: Score 0.0
Rank
Model
Score
Extra Training Data
Paper
Code
Year
1
Best Learner
0
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
1
Full Tree
0
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
3
IQN
-9289
✓
Implicit Quantile Networks for Distributional Reinforcement Learning
opendilab/DI-engine
·
chainer/chainerrl
·
Kchu/DeepRL_CK
·
+16
2018
4
MuZero
-29968.36
✓
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
werner-duvaud/muzero-general
·
opendilab/LightZero
·
koulanurag/muzero-pytorch
·
+15
2019
5
R2D2
-30021.7
✓
Recurrent Experience Replay in Distributed Reinforcement Learning
opendilab/DI-engine
·
michaelnny/deep_rl_zoo
·
garymm/earl
2019
6
IMPALA (deep)
-10180.38
✓
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
ray-project/ray
·
opendilab/DI-engine
·
deepmind/haiku
·
+21
2018
7
FQF
-9085.3
✓
Fully Parameterized Quantile Function for Distributional Reinforcement Learning
opendilab/DI-engine
·
ku2482/fqf-iqn-qrdqn.pytorch
·
ku2482/rljax
·
+3
2019
8
NoisyNet-Dueling
-7550
✓
Noisy Networks for Exploration
opendilab/DI-engine
·
Curt-Park/rainbow-is-all-you-need
·
chainer/chainerrl
·
+12
2017
9
CGP
-9011
✓
Evolving simple programs for playing Atari games
ShuhuaGao/gpFlappyBird
·
JacobLaney/cgp-tetris
2018
10
QR-DQN-1
-9324
✓
Distributional Reinforcement Learning with Quantile Regression
DLR-RM/stable-baselines3
·
facebookresearch/Horizon
·
facebookresearch/ReAgent
·
+14
2017
11
Ape-X
-10789.9
✓
Distributed Prioritized Experience Replay
ray-project/ray
·
vwxyzjn/cleanrl
·
opendilab/DI-engine
·
+12
2018
12
Advantage Learning
-13264.51
✓
Increasing the Action Gap: New Operators for Reinforcement Learning
janhuenermann/neurojs
·
chainer/chainerrl
2015
13
Go-Explore
-3660
✓
First return, then explore
uber-research/go-explore
·
qgallouedec/lge
2020
14
Agent57
-4202.6
✓
Agent57: Outperforming the Atari Human Benchmark
michaelnny/deep_rl_zoo
·
pocokhc/agent57
·
yuta0821/agent57_pytorch
·
+2
2020
15
DreamerV2
-9299
✓
Mastering Atari with Discrete World Models
opendilab/DI-engine
·
danijar/dreamerv2
·
andrejorsula/drl_grasping
·
+6
2020
16
Recurrent Rational DQN Average
-23582
✓
Adaptive Rational Activations to Boost Deep Reinforcement Learning
ml-research/rational_activations
·
ml-research/rational_rl
·
k4ntz/activation-functions
·
+1
2021
17
Rational DQN Average
-23487
✓
Adaptive Rational Activations to Boost Deep Reinforcement Learning
ml-research/rational_activations
·
ml-research/rational_rl
·
k4ntz/activation-functions
·
+1
2021
18
MuZero (Res2 Adam)
-30000
✓
Online and Offline Reinforcement Learning by Planning with a Learned Model
DHDev0/Muzero-unplugged
·
enpasos/muzero
2021
19
GDI-I3
-6774
GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
2021
19
GDI-I3
-6774
Generalized Data Distribution Iteration
2022
21
GDI-H3
-6025
Generalized Data Distribution Iteration
2022
22
DNA
-29974
DNA: Proximal Policy Optimization with a Dual Network Architecture
maitchison/PPO
2022
23
ASL DDQN
-8295.4
Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity
xinjinghao/color
2023
24
Best Learner
0
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
24
Full Tree
0
The Arcade Learning Environment: An Evaluation Platform for General Agents
mgbellemare/Arcade-Learning-Environment
·
kenjyoung/MinAtar
·
nandomp/AICollaboratory
2012
26
IQN
-9289
✓
Implicit Quantile Networks for Distributional Reinforcement Learning
opendilab/DI-engine
·
chainer/chainerrl
·
Kchu/DeepRL_CK
·
+16
2018
27
MuZero
-29968.36
✓
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
werner-duvaud/muzero-general
·
opendilab/LightZero
·
koulanurag/muzero-pytorch
·
+15
2019
28
R2D2
-30021.7
✓
Recurrent Experience Replay in Distributed Reinforcement Learning
opendilab/DI-engine
·
michaelnny/deep_rl_zoo
·
garymm/earl
2019
29
IMPALA (deep)
-10180.38
✓
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
ray-project/ray
·
opendilab/DI-engine
·
deepmind/haiku
·
+21
2018
30
FQF
-9085.3
✓
Fully Parameterized Quantile Function for Distributional Reinforcement Learning
opendilab/DI-engine
·
ku2482/fqf-iqn-qrdqn.pytorch
·
ku2482/rljax
·
+3
2019
31
NoisyNet-Dueling
-7550
✓
Noisy Networks for Exploration
opendilab/DI-engine
·
Curt-Park/rainbow-is-all-you-need
·
chainer/chainerrl
·
+12
2017
32
CGP
-9011
✓
Evolving simple programs for playing Atari games
ShuhuaGao/gpFlappyBird
·
JacobLaney/cgp-tetris
2018
33
QR-DQN-1
-9324
✓
Distributional Reinforcement Learning with Quantile Regression
DLR-RM/stable-baselines3
·
facebookresearch/Horizon
·
facebookresearch/ReAgent
·
+14
2017
34
Ape-X
-10789.9
✓
Distributed Prioritized Experience Replay
ray-project/ray
·
vwxyzjn/cleanrl
·
opendilab/DI-engine
·
+12
2018
35
Advantage Learning
-13264.51
✓
Increasing the Action Gap: New Operators for Reinforcement Learning
janhuenermann/neurojs
·
chainer/chainerrl
2015
36
Go-Explore
-3660
✓
First return, then explore
uber-research/go-explore
·
qgallouedec/lge
2020
37
Agent57
-4202.6
✓
Agent57: Outperforming the Atari Human Benchmark
michaelnny/deep_rl_zoo
·
pocokhc/agent57
·
yuta0821/agent57_pytorch
·
+2
2020
38
DreamerV2
-9299
✓
Mastering Atari with Discrete World Models
opendilab/DI-engine
·
danijar/dreamerv2
·
andrejorsula/drl_grasping
·
+6
2020
39
Recurrent Rational DQN Average
-23582
✓
Adaptive Rational Activations to Boost Deep Reinforcement Learning
ml-research/rational_activations
·
ml-research/rational_rl
·
k4ntz/activation-functions
·
+1
2021
40
Rational DQN Average
-23487
✓
Adaptive Rational Activations to Boost Deep Reinforcement Learning
ml-research/rational_activations
·
ml-research/rational_rl
·
k4ntz/activation-functions
·
+1
2021
41
MuZero (Res2 Adam)
-30000
✓
Online and Offline Reinforcement Learning by Planning with a Learned Model
DHDev0/Muzero-unplugged
·
enpasos/muzero
2021
42
GDI-I3
-6774
GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
2021
42
GDI-I3
-6774
Generalized Data Distribution Iteration
2022
44
GDI-H3
-6025
Generalized Data Distribution Iteration
2022
45
DNA
-29974
DNA: Proximal Policy Optimization with a Dual Network Architecture
maitchison/PPO
2022
46
ASL DDQN
-8295.4
Train a Real-world Local Path Planner in One Hour via Partially Decoupled Reinforcement Learning and Vectorized Diversity
xinjinghao/color
2023
1–46 / 46
페이지당
10
20
50
100