paper-with-me

홈 › Papers

Parallel bandit architecture based on laser chaos for reinforcement learning

2022-05-19 · Takashi Urushibara, Nicolas Chauvet, Satoshi Kochi, Satoshi Sunada, Kazutaka Kanno, Atsushi Uchida, Ryoichi Horisaki, Makoto Naruse

Accelerating artificial intelligence by photonics is an active field of study aiming to exploit the unique properties of photons. Reinforcement learning is an important branch of machine learning, and photonic decision-making principles have been demonstrated with respect to the multi-armed bandit problems. However, reinforcement learning could involve a massive number of states, unlike previously demonstrated bandit problems where the number of states is only one. Q-learning is a well-known approach in reinforcement learning that can deal with many states. The architecture of Q-learning, however, does not fit well photonic implementations due to its separation of update rule and the action selection. In this study, we organize a new architecture for multi-state reinforcement learning as a parallel array of bandit problems in order to benefit from photonic decision-makers, which we call parallel bandit architecture for reinforcement learning or PBRL in short. Taking a cart-pole balancing problem as an instance, we demonstrate that PBRL adapts to the environment in fewer time steps than Q-learning. Furthermore, PBRL yields faster adaptation when operated with a chaotic laser time series than the case with uniformly distributed pseudorandom numbers where the autocorrelation inherent in the laser chaos provides a positive effect. We also find that the variety of states that the system undergoes during the learning phase exhibits completely different properties between PBRL and Q-learning. The insights obtained through the present study are also beneficial for existing computing platforms, not just photonic realizations, in accelerating performances by the PBRL algorithms and correlated random sequences.

📄 PDF Abstract BibTeX arXiv:2205.09543

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingQ-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Time Series Analysis

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Scalable photonic reinforcement learning by time-division multiplexing of laser chaos

2018-03-26 · Makoto Naruse, Takatomo Mihana, Hirokazu Hori, Hayato Saigo 외

Reinforcement learning involves decision making in dynamic and uncertain environments and constitutes a crucial element of artificial intelligence. In our previous work, we experimentally demonstrated that the ultrafast …

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

Ultrafast photonic reinforcement learning based on laser chaos

2017-04-14 · Makoto Naruse, Yuta Terashima, Atsushi Uchida, Song-Ju Kim

Reinforcement learning involves decision making in dynamic and uncertain environments, and constitutes one important element of artificial intelligence (AI). In this paper, we experimentally demonstrate that the ultrafas…

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Parallel photonic accelerator for decision making using optical spatiotemporal chaos

2022-10-12 · Kensei Morijiri, Kento Takehana, Takatomo Mihana, Kazutaka Kanno 외

Photonic accelerators have attracted increasing attention in artificial intelligence applications. The multi-armed bandit problem is a fundamental problem of decision making using reinforcement learning. However, the sca…

Decision Making

Arm order recognition in multi-armed bandit problem with laser chaos time series

2020-05-26 · Naoki Narisawa, Nicolas Chauvet, Mikio Hasegawa, Makoto Naruse

By exploiting ultrafast and irregular time series generated by lasers with delayed feedback, we have previously demonstrated a scalable algorithm to solve multi-armed bandit (MAB) problems utilizing the time-division mul…

Irregular Time SeriesTime SeriesTime Series Analysis

Conflict-free joint decision by lag and zero-lag synchronization in laser network

2023-07-28 · Hisako Ito, Takatomo Mihana, Ryoichi Horisaki, Makoto Naruse

With the end of Moore's Law and the increasing demand for computing, photonic accelerators are garnering considerable attention. This is due to the physical characteristics of light, such as high bandwidth and multiplici…

Collision AvoidanceDecision Making