paper-with-me

홈 › Papers

Relative gradient optimization of the Jacobian term in unsupervised deep learning

2020-06-26 · NeurIPS 2020 12 · Luigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf, Aapo Hyvärinen

Learning expressive probabilistic models correctly describing the data is a ubiquitous problem in machine learning. A popular approach for solving it is mapping the observations into a representation space with a simple joint distribution, which can typically be written as a product of its marginals -- thus drawing a connection with the field of nonlinear independent component analysis. Deep density models have been widely used for this task, but their maximum likelihood based training requires estimating the log-determinant of the Jacobian and is computationally expensive, thus imposing a trade-off between computation and expressive power. In this work, we propose a new approach for exact training of such neural networks. Based on relative gradients, we exploit the matrix structure of neural network parameters to compute updates efficiently even in high-dimensional spaces; the computational cost of the training is quadratic in the input size, in contrast with the cubic scaling of naive approaches. This allows fast training with objective functions involving the log-determinant of the Jacobian, without imposing constraints on its structure, in stark contrast to autoregressive normalizing flows.

📄 PDF Abstract BibTeX arXiv:2006.15090

Code (1)

fissoreg/relative-gradient-jacobian 공식 구현 jax

Similar Papers 제목 키워드 기반

JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models

2026-07-20 · Ruiyi Ding, Jie Li, He Kang, Ziyan Liu 외 arxiv

Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushingl…

Reinforcement Learning

Jacobian Descent for Multi-Objective Optimization

2024-06-23 · Pierre Quinton, Valérian Rey

Many optimization problems require balancing multiple conflicting objectives. As gradient descent is limited to single-objective optimization, we introduce its direct generalization: Jacobian descent (JD). This algorithm…

image-classificationImage ClassificationMulti-Task Learning

The Curse of Unrolling: Rate of Differentiating Through Optimization

2022-09-27 · Damien Scieur, Quentin Bertrand, Gauthier Gidel, Fabian Pedregosa

Computing the Jacobian of the solution of an optimization problem is a central problem in machine learning, with applications in hyperparameter optimization, meta-learning, optimization as a layer, and dataset distillati…

Dataset DistillationHyperparameter OptimizationMeta-LearningRolling Shutter Correction

Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence

2024-10-03 · Ruichen Jiang, Aryan Mokhtari

In this paper, we propose a quasi-Newton method for solving smooth and monotone nonlinear equations, including unconstrained minimization and minimax optimization as special cases. For the strongly monotone setting, we e…

RecurJac: An Efficient Recursive Algorithm for Bounding Jacobian Matrix of Neural Networks and Its Applications

2018-10-28 · Huan Zhang, Pengchuan Zhang, Cho-Jui Hsieh

The Jacobian matrix (or the gradient for single-output networks) is directly related to many important properties of neural networks, such as the function landscape, stationary points, (local) Lipschitz constants and rob…