paper-with-me

홈 › Papers

Virtual Action Actor-Critic Framework for Exploration (Student Abstract)

2023-11-06 · Bumgeun Park, TaeYoung Kim, Quoc-Vinh Lai-Dang, Dongsoo Har

Efficient exploration for an agent is challenging in reinforcement learning (RL). In this paper, a novel actor-critic framework namely virtual action actor-critic (VAAC), is proposed to address the challenge of efficient exploration in RL. This work is inspired by humans' ability to imagine the potential outcomes of their actions without actually taking them. In order to emulate this ability, VAAC introduces a new actor called virtual actor (VA), alongside the conventional actor-critic framework. Unlike the conventional actor, the VA takes the virtual action to anticipate the next state without interacting with the environment. With the virtual policy following a Gaussian distribution, the VA is trained to maximize the anticipated novelty of the subsequent state resulting from a virtual action. If any next state resulting from available actions does not exhibit high anticipated novelty, training the VA leads to an increase in the virtual policy entropy. Hence, high virtual policy entropy represents that there is no room for exploration. The proposed VAAC aims to maximize a modified Q function, which combines cumulative rewards and the negative sum of virtual policy entropy. Experimental results show that the VAAC improves the exploration performance compared to existing algorithms.

📄 PDF Abstract BibTeX arXiv:2311.02916

Code (0)

등록된 구현이 없습니다.

Tasks

Efficient ExplorationReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

SELU: Self-Learning Embodied MLLMs in Unknown Environments

2024-10-04 · Boyu Li, Haobin Jiang, Ziluo Ding, Xinrun Xu 외

Recently, multimodal large language models (MLLMs) have demonstrated strong visual understanding and decision-making capabilities, enabling the exploration of autonomously improving MLLMs in unknown environments. However…

Decision MakingSelf-Learning

STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft

2024-06-17 · Zhonghan Zhao, Wenhao Chai, Xuan Wang, Ke Ma 외

Building an embodied agent system with a large language model (LLM) as its core is a promising direction. Due to the significant costs and uncontrollable factors associated with deploying and training such agents in the …

Knowledge DistillationLanguage ModelingLanguage ModellingLarge Language Model+1

TAAC: Temporally Abstract Actor-Critic for Continuous Control

2021-04-13 · NeurIPS 2021 12 · Haonan Yu, Wei Xu, Haichao Zhang

We present temporally abstract actor-critic (TAAC), a simple but effective off-policy RL algorithm that incorporates closed-loop temporal abstraction into the actor-critic framework. TAAC adds a second-stage binary polic…

continuous-controlContinuous Control

Improving Exploration in Soft-Actor-Critic with Normalizing Flows Policies

2019-06-06 · Patrick Nadeem Ward, Ariella Smofsky, Avishek Joey Bose

Deep Reinforcement Learning (DRL) algorithms for continuous action spaces are known to be brittle toward hyperparameters as well as \cut{being}sample inefficient. Soft Actor Critic (SAC) proposes an off-policy deep actor…

Deep Reinforcement LearningReinforcement LearningReinforcement Learning (RL)

Better Exploration with Optimistic Actor-Critic

2019-10-28 · Kamil Ciosek, Quan Vuong, Robert Loftin, Katja Hofmann

Actor-critic methods, a type of model-free Reinforcement Learning, have been successfully applied to challenging tasks in continuous control, often achieving state-of-the art performance. However, wide-scale adoption of …

continuous-controlContinuous ControlEfficient ExplorationReinforcement Learning