paper-with-me

홈 › Papers

Which Experiences Are Influential for Your Agent? Policy Iteration with Turn-over Dropout

2023-01-26 · Takuya Hiraoka, Takashi Onishi, Yoshimasa Tsuruoka

In reinforcement learning (RL) with experience replay, experiences stored in a replay buffer influence the RL agent's performance. Information about the influence is valuable for various purposes, including experience cleansing and analysis. One method for estimating the influence of individual experiences is agent comparison, but it is prohibitively expensive when there is a large number of experiences. In this paper, we present PI+ToD as a method for efficiently estimating the influence of experiences. PI+ToD is a policy iteration that efficiently estimates the influence of experiences by utilizing turn-over dropout. We demonstrate the efficiency of PI+ToD with experiments in MuJoCo environments.

📄 PDF Abstract BibTeX arXiv:2301.11168

Code (1)

takuyahiraoka/which-experiences-are-influential-for-your-agent 공식 구현 pytorch

Tasks

MuJoCoreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Which Experiences Are Influential for RL Agents? Efficiently Estimating The Influence of Experiences

2024-05-23 · Takuya Hiraoka, Guanquan Wang, Takashi Onishi, Yoshimasa Tsuruoka

In reinforcement learning (RL) with experience replay, experiences stored in a replay buffer influence the RL agent's performance. Information about how these experiences influence the agent's performance is valuable for…

Reinforcement Learning (RL)

Experience Replay Optimization

2019-06-19 · Daochen Zha, Kwei-Herng Lai, Kaixiong Zhou, Xia Hu

Experience replay enables reinforcement learning agents to memorize and reuse past experiences, just as humans replay memories for the situation at hand. Contemporary off-policy algorithms either replay past experiences …

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1

Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy

2020-09-29 · Yunshu Du, Garrett Warnell, Assefaw Gebremedhin, Peter Stone 외

Experience replay (ER) improves the data efficiency of off-policy reinforcement learning (RL) algorithms by allowing an agent to store and reuse its past experiences in a replay buffer. While many techniques have been pr…

Atari GamesReinforcement Learning (RL)

Depending on yourself when you should: Mentoring LLM with RL agents to become the master in cybersecurity games

2024-03-26 · Yikuan Yan, Yaolun Zhang, Keman Huang

Integrating LLM and reinforcement learning (RL) agent effectively to achieve complementary performance is critical in high stake tasks like cybersecurity operations. In this study, we introduce SecurityBot, a LLM agent m…

Reinforcement Learning (RL)

Plan Your Target and Learn Your Skills: State-Only Imitation Learning via Decoupled Policy Optimization

2021-09-29 · NeurIPS 2021 12 · Minghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang 외

State-only imitation learning (SOIL) enables agents to learn from massive demonstrations without explicit action or reward information. However, previous methods attempt to learn the implicit state-to-action mapping poli…

Imitation LearningReinforcement Learning (RL)