An Adiabatic Theorem for Policy Tracking with TD-learning
We evaluate the ability of temporal difference learning to track the reward function of a policy as it changes over time. Our results apply a new adiabatic theorem that bounds the mixing time of time-inhomogeneous Markov chains. We derive finite-time bounds for tabular temporal difference learning and $Q$-learning when the policy used for training changes in time. To achieve this, we develop bounds for stochastic approximation under asynchronous adiabatic updates.
Code (0)
등록된 구현이 없습니다.
Tasks
Q-LearningSimilar Papers 제목 키워드 기반
The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport
We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family of Schrödinger operators we call Score Hamiltonians, built from the learned scor…
Adiabatic Quantum Computing for Multi Object Tracking
Multi-Object Tracking (MOT) is most often approached in the tracking-by-detection paradigm, where object detections are associated through time. The association step naturally leads to discrete optimization problems. As …
Multi-Object TrackingObjectObject TrackingWarm Starts, Cold States: Exploiting Adiabaticity for Variational Ground-States
Reliable preparation of many-body ground states is an essential task in quantum computing, with applications spanning areas from chemistry and materials modeling to quantum optimization and benchmarking. A variety of app…
Quantum Adiabatic Algorithm Design using Reinforcement Learning
Quantum algorithm design plays a crucial role in exploiting the computational advantage of quantum devices. Here we develop a deep-reinforcement-learning based approach for quantum adiabatic algorithm design. Our approac…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Machine-learning parameter tracking with partial state observation
Complex and nonlinear dynamical systems often involve parameters that change with time, accurate tracking of which is essential to tasks such as state estimation, prediction, and control. Existing machine-learning method…
State Estimation