Which Experiences Are Influential for Your Agent? Policy Iteration with Turn-over Dropout
In reinforcement learning (RL) with experience replay, experiences stored in a replay buffer influence the RL agent's performance. Information about the influence is valuable for various purposes, including experience cleansing and analysis. One method for estimating the influence of individual experiences is agent comparison, but it is prohibitively expensive when there is a large number of experiences. In this paper, we present PI+ToD as a method for efficiently estimating the influence of experiences. PI+ToD is a policy iteration that efficiently estimates the influence of experiences by utilizing turn-over dropout. We demonstrate the efficiency of PI+ToD with experiments in MuJoCo environments.
Code (1)
Tasks
MuJoCoreinforcement-learningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Which Experiences Are Influential for RL Agents? Efficiently Estimating The Influence of Experiences
In reinforcement learning (RL) with experience replay, experiences stored in a replay buffer influence the RL agent's performance. Information about how these experiences influence the agent's performance is valuable for…
Reinforcement Learning (RL)Experience Replay Optimization
Experience replay enables reinforcement learning agents to memorize and reuse past experiences, just as humans replay memories for the situation at hand. Contemporary off-policy algorithms either replay past experiences …
continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy
Experience replay (ER) improves the data efficiency of off-policy reinforcement learning (RL) algorithms by allowing an agent to store and reuse its past experiences in a replay buffer. While many techniques have been pr…
Atari GamesReinforcement Learning (RL)Depending on yourself when you should: Mentoring LLM with RL agents to become the master in cybersecurity games
Integrating LLM and reinforcement learning (RL) agent effectively to achieve complementary performance is critical in high stake tasks like cybersecurity operations. In this study, we introduce SecurityBot, a LLM agent m…
Reinforcement Learning (RL)Plan Your Target and Learn Your Skills: State-Only Imitation Learning via Decoupled Policy Optimization
State-only imitation learning (SOIL) enables agents to learn from massive demonstrations without explicit action or reward information. However, previous methods attempt to learn the implicit state-to-action mapping poli…
Imitation LearningReinforcement Learning (RL)