paper-with-me

홈 › Papers

Streaming Linear System Identification with Reverse Experience Replay

2021-03-10 · NeurIPS 2021 12 · Prateek Jain, Suhas S Kowshik, Dheeraj Nagaraj, Praneeth Netrapalli

We consider the problem of estimating a linear time-invariant (LTI) dynamical system from a single trajectory via streaming algorithms, which is encountered in several applications including reinforcement learning (RL) and time-series analysis. While the LTI system estimation problem is well-studied in the {\em offline} setting, the practically important streaming/online setting has received little attention. Standard streaming methods like stochastic gradient descent (SGD) are unlikely to work since streaming points can be highly correlated. In this work, we propose a novel streaming algorithm, SGD with Reverse Experience Replay ($\mathsf{SGD}-\mathsf{RER}$), that is inspired by the experience replay (ER) technique popular in the RL literature. $\mathsf{SGD}-\mathsf{RER}$ divides data into small buffers and runs SGD backwards on the data stored in the individual buffers. We show that this algorithm exactly deconstructs the dependency structure and obtains information theoretically optimal guarantees for both parameter error and prediction error. Thus, we provide the first -- to the best of our knowledge -- optimal SGD-style algorithm for the classical problem of linear system identification with a first order oracle. Furthermore, $\mathsf{SGD}-\mathsf{RER}$ can be applied to more general settings like sparse LTI identification with known sparsity pattern, and non-linear dynamical systems. Our work demonstrates that the knowledge of data dependency structure can aid us in designing statistically and computationally efficient algorithms which can "decorrelate" streaming samples.

📄 PDF Abstract BibTeX arXiv:2103.05896

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)Time Series Analysis

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…
Experience Replay Experience Replay is a replay memory technique used in reinforcement learning where we store the agent’s experiences at each time-step, $e\_{t} = \left(s\_{t}, a\_{t}, r\_{t},…

Similar Papers 제목 키워드 기반

Distributed Online System Identification for LTI Systems Using Reverse Experience Replay

2022-07-03 · Ting-Jui Chang, Shahin Shahrampour

Identification of linear time-invariant (LTI) systems plays an important role in control and reinforcement learning. Both asymptotic and finite-time offline system identification are well-studied in the literature. For o…

Near-optimal Offline and Streaming Algorithms for Learning Non-Linear Dynamical Systems

2021-05-24 · NeurIPS 2021 12 · Prateek Jain, Suhas S Kowshik, Dheeraj Nagaraj, Praneeth Netrapalli

We consider the setting of vector valued non-linear dynamical systems $X_{t+1} = \phi(A^* X_t) + \eta_t$, where $\eta_t$ is unbiased noise and $\phi : \mathbb{R} \to \mathbb{R}$ is a known link function that satisfies ce…

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

2026-05-28 · Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu 외 arxiv

Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements for temporal consist…

Reverse Experience Replay

2019-10-19 · Egor Rotinov

This paper describes an improvement in Deep Q-learning called Reverse Experience Replay (also RER) that solves the problem of sparse rewards and helps to deal with reward maximizing tasks by sampling transitions successi…

Q-Learning

Online Weak-form Sparse Identification of Partial Differential Equations

2022-03-08 · Daniel A. Messenger, Emiliano Dall'Anese, David M. Bortz

This paper presents an online algorithm for identification of partial differential equations (PDEs) based on the weak-form sparse identification of nonlinear dynamics algorithm (WSINDy). The algorithm is online in a sens…

Form