Montezuma's Revenge
1개 벤치마크 · 논문 64편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Rainbow: Combining Improvements in Deep Reinforcement Learning
Exploration by Random Network Distillation
Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
Go-Explore: a New Approach for Hard-Exploration Problems
A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs
Reinforcement Learning with Latent Flow
Papers
Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games
World-model synthesis aims to turn interaction experience into an internal model of environment dynamics. Existing symbolic approaches often fit observed transitions or mixtures of local rules, but they do not produce a …
Montezuma's RevengeDecoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration
The process of discovery requires active exploration -- the act of collecting new and informative data. However, efficient autonomous exploration remains a major unsolved problem. The dominant paradigm addresses this cha…
Reinforcement LearningMontezuma's RevengeLLM-assisted Semantic Option Discovery for Facilitating Adaptive Deep Reinforcement Learning
Despite achieving remarkable success in complex tasks, Deep Reinforcement Learning (DRL) is still suffering from critical issues in practical applications, such as low data efficiency, lack of interpretability, and limit…
Reinforcement LearningMontezuma's RevengeGeneral KnowledgeAction-Dependent Optimality-Preserving Reward Shaping
Recent RL research has utilized reward shaping--particularly complex shaping rewards such as intrinsic motivation (IM)--to encourage agent exploration in sparse-reward environments. While often effective, ``reward hackin…
Montezuma's RevengePoE-World: Compositional World Modeling with Products of Programmatic Experts
Learning how the world works is central to building AI agents that can adapt to complex environments. Traditional world models based on deep learning demand vast amounts of training data, and do not flexibly update their…
Montezuma's RevengeProgram SynthesisA Study of Plasticity Loss in On-Policy Deep Reinforcement Learning
Continual learning with deep neural networks presents challenges distinct from both the fixed-dataset and convex continual learning regimes. One such challenge is plasticity loss, wherein a neural network trained in an o…
Continual LearningDeep Reinforcement LearningMontezuma's RevengeReinforcement Learning (RL)