paper-with-me

Montezuma's Revenge

1개 벤치마크 · 논문 64편 · 이 태스크의 논문 보기 →

Benchmarks

Most implemented

Exploration by Random Network Distillation

2018-10-30 · 구현 22개

Reinforcement Learning with Latent Flow

2021-01-06 · 구현 2개

Papers

Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games

2026-06-14 · Yifei Dong, Mingen Zheng, Linquan Wu, Jeff Z. Pan 외 arxiv

World-model synthesis aims to turn interaction experience into an internal model of environment dynamics. Existing symbolic approaches often fit observed transitions or mixtures of local rules, but they do not produce a …

Montezuma's Revenge

Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration

2026-03-23 · Zakaria Mhammedi, James Cohan arxiv

The process of discovery requires active exploration -- the act of collecting new and informative data. However, efficient autonomous exploration remains a major unsolved problem. The dominant paradigm addresses this cha…

Reinforcement LearningMontezuma's Revenge

LLM-assisted Semantic Option Discovery for Facilitating Adaptive Deep Reinforcement Learning

2026-03-02 · Chang Yao, Jinghui Qin, Kebing Jin, Hankz Hankui Zhuo arxiv

Despite achieving remarkable success in complex tasks, Deep Reinforcement Learning (DRL) is still suffering from critical issues in practical applications, such as low data efficiency, lack of interpretability, and limit…

Reinforcement LearningMontezuma's RevengeGeneral Knowledge

Action-Dependent Optimality-Preserving Reward Shaping

2025-05-19 · Grant C. Forbes, JianXun Wang, Leonardo Villalobos-Arias, Arnav Jhala 외

Recent RL research has utilized reward shaping--particularly complex shaping rewards such as intrinsic motivation (IM)--to encourage agent exploration in sparse-reward environments. While often effective, ``reward hackin…

Montezuma's Revenge

PoE-World: Compositional World Modeling with Products of Programmatic Experts

2025-05-16 · Wasu Top Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller 외

Learning how the world works is central to building AI agents that can adapt to complex environments. Traditional world models based on deep learning demand vast amounts of training data, and do not flexibly update their…

Montezuma's RevengeProgram Synthesis

A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning

2024-05-29 · Arthur Juliani, Jordan T. Ash

Continual learning with deep neural networks presents challenges distinct from both the fixed-dataset and convex continual learning regimes. One such challenge is plasticity loss, wherein a neural network trained in an o…

Continual LearningDeep Reinforcement LearningMontezuma's RevengeReinforcement Learning (RL)

전체 64편 보기 →