paper-with-me

홈 › Papers

GO Hessian for Expectation-Based Objectives

2020-06-16 · Yulai Cong, Miaoyun Zhao, Jianqiao Li, Junya Chen, Lawrence Carin

An unbiased low-variance gradient estimator, termed GO gradient, was proposed recently for expectation-based objectives $\mathbb{E}_{q_{\boldsymbol{\gamma}}(\boldsymbol{y})} [f(\boldsymbol{y})]$, where the random variable (RV) $\boldsymbol{y}$ may be drawn from a stochastic computation graph with continuous (non-reparameterizable) internal nodes and continuous/discrete leaves. Upgrading the GO gradient, we present for $\mathbb{E}_{q_{\boldsymbol{\boldsymbol{\gamma}}}(\boldsymbol{y})} [f(\boldsymbol{y})]$ an unbiased low-variance Hessian estimator, named GO Hessian. Considering practical implementation, we reveal that GO Hessian is easy-to-use with auto-differentiation and Hessian-vector products, enabling efficient cheap exploitation of curvature information over stochastic computation graphs. As representative examples, we present the GO Hessian for non-reparameterizable gamma and negative binomial RVs/nodes. Based on the GO Hessian, we design a new second-order method for $\mathbb{E}_{q_{\boldsymbol{\boldsymbol{\gamma}}}(\boldsymbol{y})} [f(\boldsymbol{y})]$, with rigorous experiments conducted to verify its effectiveness and efficiency.

📄 PDF Abstract BibTeX arXiv:2006.08873

Code (1)

YulaiCong/GOHessian 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Convergence Analysis of Block Coordinate Algorithms with Determinantal Sampling

2019-10-25 · Mojmír Mutný, Michał Dereziński, Andreas Krause

We analyze the convergence rate of the randomized Newton-like method introduced by Qu et. al. (2016) for smooth and convex objectives, which uses random coordinate blocks of a Hessian-over-approximation matrix $\bM$ inst…

Point Processes

Fast Unconstrained Optimization via Hessian Averaging and Adaptive Gradient Sampling Methods

2024-08-14 · Thomas O'Leary-Roseberry, Raghu Bollapragada

We consider minimizing finite-sum and expectation objective functions via Hessian-averaging based subsampled Newton methods. These methods allow for gradient inexactness and have fixed per-iteration Hessian approximation…

Efficient Learning of Restricted Boltzmann Machines Using Covariance Estimates

2018-10-25 · Vidyadhar Upadhya, P. S. Sastry

Learning RBMs using standard algorithms such as CD(k) involves gradient descent on the negative log-likelihood. One of the terms in the gradient, which involves expectation w.r.t. the model distribution, is intractable a…

A Unified Theory of $θ$-Expectations

2025-07-27 · Qian Qi arxiv

We derive a new class of non-linear expectations from first-principles deterministic chaotic dynamics. The homogenization of the system's skew-adjoint microscopic generator is achieved using the spectral theory of transf…

Efficient Score Computation and Expectation-Maximization Algorithm in Regime-Switching Models

2022-05-03 · Chaojun Li, Shi Qiu

This study proposes an efficient algorithm for score computation for regime-switching models, and derived from which, an efficient expectation-maximization (EM) algorithm. Different from existing algorithms, this algorit…