paper-with-me

홈 › Papers

On the Variance of Unbiased Online Recurrent Optimization

2019-02-06 · Tim Cooijmans, James Martens

The recently proposed Unbiased Online Recurrent Optimization algorithm (UORO, arXiv:1702.05043) uses an unbiased approximation of RTRL to achieve fully online gradient-based learning in RNNs. In this work we analyze the variance of the gradient estimate computed by UORO, and propose several possible changes to the method which reduce this variance both in theory and practice. We also contribute significantly to the theoretical and intuitive understanding of UORO (and its existing variance reduction technique), and demonstrate a fundamental connection between its gradient estimate and the one that would be computed by REINFORCE if small amounts of noise were added to the RNN's hidden units.

📄 PDF Abstract BibTeX arXiv:1902.02405

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

UORO 설명 없음

Similar Papers 제목 키워드 기반

Unbiased Online Recurrent Optimization

2017-02-16 · ICLR 2018 1 · Corentin Tallec, Yann Ollivier

The novel Unbiased Online Recurrent Optimization (UORO) algorithm allows for online learning of general recurrent computational graphs such as recurrent network models. It works in a streaming fashion and avoids backtrac…

Enhanced Federated Optimization: Adaptive Unbiased Client Sampling with Reduced Variance

2023-10-04 · Dun Zeng, Zenglin Xu, Yu Pan, Xu Luo 외

Federated Learning (FL) is a distributed learning paradigm to train a global model across multiple devices without collecting local data. In FL, a server typically selects a subset of clients for each training round to o…

Federated Learning

Low-Variance Gradient Estimation in Unrolled Computation Graphs with ES-Single

2023-04-21 · Paul Vicol, Zico Kolter, Kevin Swersky

We propose an evolution strategies-based algorithm for estimating gradients in unrolled computation graphs, called ES-Single. Similarly to the recently-proposed Persistent Evolution Strategies (PES), ES-Single is unbiase…

Hyperparameter Optimization

Bi-Level Decision-Focused Causal Learning for Large-Scale Marketing Optimization: Bridging Observational and Experimental Data

2025-10-22 · Shuli Zhang, Hao Zhou, Jiaqi Zheng, Guibin Jiang 외 arxiv

Online Internet platforms require sophisticated marketing strategies to optimize user retention and platform revenue -- a classical resource allocation problem. Traditional solutions adopt a two-stage pipeline: machine l…

A Practical Sparse Approximation for Real Time Recurrent Learning

2020-06-12 · Jacob Menick, Erich Elsen, Utku Evci, Simon Osindero 외

Current methods for training recurrent neural networks are based on backpropagation through time, which requires storing a complete history of network states, and prohibits updating the weights `online' (after every time…