paper-with-me

Papers

Playing hard exploration games by watching YouTube

2018-05-29 · NeurIPS 2018 12 · Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, Nando de Freitas

Deep reinforcement learning methods traditionally struggle with tasks where environment rewards are particularly sparse. One successful method of guiding exploration in these domains is to imitate trajectories provided by a human demonstrator. However, these demonstrations are typically collected under artificial conditions, i.e. with access to the agent's exact environment setup and the demonstrator's action and reward trajectories. Here we propose a two-stage method that overcomes these limitations by relying on noisy, unaligned footage without access to such data. First, we learn to map unaligned videos from multiple sources to a common representation using self-supervised objectives constructed over both time and modality (i.e. vision and sound). Second, we embed a single YouTube video in this representation to construct a reward function that encourages an agent to imitate human gameplay. This method of one-shot imitation allows our agent to convincingly exceed human-level performance on the infamously hard exploration games Montezuma's Revenge, Pitfall! and Private Eye for the first time, even if the agent is not presented with any environment rewards.

📄 PDF Abstract BibTeX arXiv:1805.11592

Code (1)

MaxSobolMark/HardRLWithYoutube tf

Tasks

Deep Reinforcement LearningMontezuma's RevengeReinforcement Learning

Similar Papers 제목 키워드 기반

Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

2022-06-23 · Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 외

Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision dom…

Imitation LearningMinecraftreinforcement-learningReinforcement Learning (RL)

On Bonus Based Exploration Methods In The Arcade Learning Environment

2020-01-01 · ICLR 2020 1 · Adrien Ali Taiga, William Fedus, Marlos C. Machado, Aaron Courville 외

Research on exploration in reinforcement learning, as applied to Atari 2600 game-playing, has emphasized tackling difficult exploration problems such as Montezuma's Revenge (Bellemare et al., 2016). Recently, bonus-based…

Atari GamesMontezuma's RevengeReinforcement Learning

On Bonus-Based Exploration Methods in the Arcade Learning Environment

2021-09-22 · Adrien Ali Taïga, William Fedus, Marlos C. Machado, Aaron Courville 외

Research on exploration in reinforcement learning, as applied to Atari 2600 game-playing, has emphasized tackling difficult exploration problems such as Montezuma's Revenge (Bellemare et al., 2016). Recently, bonus-based…

Atari GamesMontezuma's Revenge

Inducing game rules from varying quality game play

2020-08-04 · Alastair Flynn

General Game Playing (GGP) is a framework in which an artificial intelligence program is required to play a variety of games successfully. It acts as a test bed for AI and motivator of research. The AI is given a random …

Inductive logic programming

Computer-Generated Music for Tabletop Role-Playing Games

2020-08-16 · Lucas N. Ferreira, Levi H. S. Lelis, Jim Whitehead

In this paper we present Bardo Composer, a system to generate background music for tabletop role-playing games. Bardo Composer uses a speech recognition system to translate player speech into text, which is classified ac…

speech-recognitionSpeech Recognition