Twin actor twin delayed deep deterministic policy gradient (TATD3) learning for batch process control
Control of batch processes is a difficult task due to their complex nonlinear dynamics and unsteady-state operating conditions within batch and batch-to-batch. It is expected that some of these challenges can be addressed by developing control strategies that directly interact with the process and learning from experiences. Recent studies in the literature have indicated the advantage of having an ensemble of actors in actor-critic Reinforcement Learning (RL) frameworks for improving the policy. The present study proposes an actor-critic RL algorithm, namely, twin actor twin delayed deep deterministic policy gradient (TATD3), by incorporating twin actor networks in the existing twin-delayed deep deterministic policy gradient (TD3) algorithm for the continuous control. In addition, two types of novel reward functions are also proposed for TATD3 controller. We showcase the efficacy of the TATD3 based controller for various batch process examples by comparing it with some of the existing RL algorithms presented in the literature.
Code (0)
등록된 구현이 없습니다.
Tasks
continuous-controlContinuous ControlReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DATD3: Depthwise Attention Twin Delayed Deep Deterministic Policy Gradient For Model Free Reinforcement Learning Under Output Feedback Control
Reinforcement learning in real-world applications often involves output-feedback settings, where the agent receives only partial state information. To address this challenge, we propose the Output-Feedback Markov Decisio…
continuous-controlContinuous ControlDecision MakingTT-DAC-PS: Twin-Target Deterministic Actor-Critic with Policy Smoothing for Optimal Trade Execution
This study addresses the optimal execution of large stock sell programs by introducing TT-DAC-PS (Twin-Target Deterministic Actor-Critic with Policy Smoothing), a deterministic actor-critic architecture that combines twi…
Control of a Twin Rotor using Twin Delayed Deep Deterministic Policy Gradient (TD3)
This paper proposes a reinforcement learning (RL) framework for controlling and stabilizing the Twin Rotor Aerodynamic System (TRAS) at specific pitch and azimuth angles and tracking a given trajectory. The complex dynam…
Reinforcement LearningWhen Do Drivers Concentrate? Attention-based Driver Behavior Modeling With Deep Reinforcement Learning
Driver distraction a significant risk to driving safety. Apart from spatial domain, research on temporal inattention is also necessary. This paper aims to figure out the pattern of drivers' temporal attention allocation.…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Dynamic Entropy Tuning in Reinforcement Learning Low-Level Quadcopter Control: Stochasticity vs Determinism
This paper explores the impact of dynamic entropy tuning in Reinforcement Learning (RL) algorithms that train a stochastic policy. Its performance is compared against algorithms that train a deterministic one. Stochastic…
Reinforcement Learning