paper-with-me

홈 › Papers

Automated Curriculum Learning by Rewarding Temporally Rare Events

2018-03-19 · Niels Justesen, Sebastian Risi

Reward shaping allows reinforcement learning (RL) agents to accelerate learning by receiving additional reward signals. However, these signals can be difficult to design manually, especially for complex RL tasks. We propose a simple and general approach that determines the reward of pre-defined events by their rarity alone. Here events become less rewarding as they are experienced more often, which encourages the agent to continually explore new types of events as it learns. The adaptiveness of this reward function results in a form of automated curriculum learning that does not have to be specified by the experimenter. We demonstrate that this \emph{Rarity of Events} (RoE) approach enables the agent to succeed in challenging VizDoom scenarios without access to the extrinsic reward from the environment. Furthermore, the results demonstrate that RoE learns a more versatile policy that adapts well to critical changes in the environment. Rewarding events based on their rarity could help in many unsolved RL environments that are characterized by sparse extrinsic rewards but a plethora of known event types.

📄 PDF Abstract BibTeX arXiv:1803.07131

Code (1)

lasseuth1/blood_bowl2 pytorch

Tasks

Reinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Automated curriculum generation through setter-solver interactions

2020-05-01 · ICLR 2020 1 · Sebastien Racaniere, Andrew Lampinen, Adam Santoro, David Reichert 외

Reinforcement learning algorithms use correlations between policies and rewards to improve agent performance. But in dynamic or sparsely rewarding environments these correlations are often too small, or rewarding even…

Automated curricula through setter-solver interactions

2019-09-27 · Sebastien Racaniere, Andrew K. Lampinen, Adam Santoro, David P. Reichert 외

Reinforcement learning algorithms use correlations between policies and rewards to improve agent performance. But in dynamic or sparsely rewarding environments these correlations are often too small, or rewarding events …

Reinforcement Learning

LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models?

2026-03-25 · Iqra Ali, Talia Tseriotou, Mahmud Elahi Akhter, Yuxiang Zhou 외 arxiv

Longitudinal NLP tasks require reasoning over temporally ordered text to detect persistence and change in human behavior and opinions. However, in-context learning with large language models struggles on tasks where mode…

Zero-Shot Dependency Parsing with Worst-Case Aware Automated Curriculum Learning

2022-03-16 · ACL 2022 5 · Miryam de Lhoneux, Sheng Zhang, Anders Søgaard

Large multilingual pretrained language models such as mBERT and XLM-RoBERTa have been found to be surprisingly effective for cross-lingual transfer of syntactic parsing models (Wu and Dredze 2019), but only between relat…

Cross-Lingual TransferDependency ParsingMulti-Task Learning

Zero-Shot Dependency Parsing with Worst-Case Aware Automated Curriculum Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Large multilingual pretrained language models such as mBERT and XLM-RoBERTa have been found to be surprisingly effective for cross-lingual transfer of syntactic parsing models Wu and Dredze (2019), but only between relat…

Cross-Lingual TransferDependency ParsingMulti-Task Learning