paper-with-me

Papers

Diversity-Enriched Option-Critic

2020-11-04 · Anand Kamat, Doina Precup

Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The option-critic framework has been demonstrated to learn temporally extended actions, represented as options, end-to-end in a model-free setting. However, feasibility of option-critic remains limited due to two major challenges, multiple options adopting very similar behavior, or a shrinking set of task relevant options. These occurrences not only void the need for temporal abstraction, they also affect performance. In this paper, we tackle these problems by learning a diverse set of options. We introduce an information-theoretic intrinsic reward, which augments the task reward, as well as a novel termination objective, in order to encourage behavioral diversity in the option set. We show empirically that our proposed method is capable of learning options end-to-end on several discrete and continuous control tasks, outperforms option-critic by a wide margin. Furthermore, we show that our approach sustainably generates robust, reusable, reliable and interpretable options, in contrast to option-critic.

📄 PDF Abstract BibTeX arXiv:2011.02565

Code (1)

anandkamat05/TDEOC 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlDiversity

Similar Papers 제목 키워드 기반

Learning Diverse Options via InfoMax Termination Critic

2020-10-06 · Yuji Kanagawa, Tomoyuki Kaneko

We consider the problem of autonomously learning reusable temporally extended actions, or options, in reinforcement learning. While options can speed up transfer learning by serving as reusable building blocks, learning …

Continuous ControlDiversityreinforcement-learningReinforcement Learning (RL)+1

Wasserstein Diversity-Enriched Regularizer for Hierarchical Reinforcement Learning

2023-08-02 · Haorui Li, Jiaqi Liang, Linjing Li, Daniel Zeng

Hierarchical reinforcement learning composites subpolicies in different hierarchies to accomplish complex tasks.Automated subpolicies discovery, which does not depend on domain knowledge, is a promising approach to gener…

DiversityHierarchical Reinforcement Learningreinforcement-learningReinforcement Learning

A Unified Algorithm Framework for Unsupervised Discovery of Skills based on Determinantal Point Process

2023-09-21 · NeurIPS 2023 11

Learning rich skills under the option framework without supervision of external rewards is at the frontier of reinforcement learning research. Existing works mainly fall into two distinctive categories: variational optio…

Fidelity-Enriched Contrastive Search: Reconciling the Faithfulness-Diversity Trade-Off in Text Generation

2023-10-23 · Wei-Lin Chen, Cheng-Kuang Wu, Hsin-Hsi Chen, Chung-Chi Chen

In this paper, we address the hallucination problem commonly found in natural language generation tasks. Language models often generate fluent and convincing content but can lack consistency with the provided source, res…

Abstractive Text SummarizationDialogue GenerationDiversityHallucination+3

A random planting model

2024-09-24 · Julian Talbot, Pascal Viot, David Colliaux

The adoption of agroecological practices will be crucial to address the challenges of climate change and biodiversity loss. Such practices favor the cultivation of plants in complex mixtures with layouts differing from t…

model