QuaRL: Quantization for Fast and Environmentally Sustainable Reinforcement Learning
Deep reinforcement learning continues to show tremendous potential in achieving task-level autonomy, however, its computational and energy demands remain prohibitively high. In this paper, we tackle this problem by applying quantization to reinforcement learning. To that end, we introduce a novel Reinforcement Learning (RL) training paradigm, \textit{ActorQ}, to speed up actor-learner distributed RL training. \textit{ActorQ} leverages 8-bit quantized actors to speed up data collection without affecting learning convergence. Our quantized distributed RL training system, \textit{ActorQ}, demonstrates end-to-end speedups \blue{between 1.5 $\times$ and 5.41$\times$}, and faster convergence over full precision training on a range of tasks (Deepmind Control Suite) and different RL algorithms (D4PG, DQN). Furthermore, we compare the carbon emissions (Kgs of CO2) of \textit{ActorQ} versus standard reinforcement learning \blue{algorithms} on various tasks. Across various settings, we show that \textit{ActorQ} enables more environmentally friendly reinforcement learning by achieving \blue{carbon emission improvements between 1.9$\times$ and 3.76$\times$} compared to training RL-agents in full-precision. We believe that this is the first of many future works on enabling computationally energy-efficient and sustainable reinforcement learning. The source code is available here for the public to use: \url{https://github.com/harvard-edge/QuaRL}.
Code (1)
Tasks
Decision MakingDeep Reinforcement LearningQuantizationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Quarl: A Learning-Based Quantum Circuit Optimizer
Optimizing quantum circuits is challenging due to the very large search space of functionally equivalent circuits and the necessity of applying transformations that temporarily decrease performance to achieve a final per…
Reinforcement Learning (RL)QForce-RL: Quantized FPGA-Optimized Reinforcement Learning Compute Engine
Reinforcement Learning (RL) has outperformed other counterparts in sequential decision-making and dynamic environment control. However, FPGA deployment is significantly resource-expensive, as associated with large number…
Decision MakingQuantizationreinforcement-learningReinforcement Learning+2Optimality and sustainability of hybrid limit cycles in the pollution control problem with regime shifts
In this paper, we consider the problem of pollution control in a system that undergoes regular regime shifts. We first show that the optimal policy of pollution abatement is periodic as well, and is described by the uniq…
An Environmentally Sustainable Closed-Loop Supply Chain Network Design under Uncertainty: Application of Optimization
Newly, the rates of energy and material consumption to augment industrial pro-duction are substantially high, thus the environmentally sustainable industrial de-velopment has emerged as the main issue of either developed…
ManagementThe Effects of Hofstede's Cultural Dimensions on Pro-Environmental Behaviour: How Culture Influences Environmentally Conscious Behaviour
The need for a more sustainable lifestyle is a key focus for several countries. Using a questionnaire survey conducted in Hungary, this paper examines how culture influences environmentally conscious behaviour. Having in…
Cultural Vocal Bursts Intensity Prediction