Energy Discrepancies: A Score-Independent Loss for Energy-Based Models
Energy-based models are a simple yet powerful class of probabilistic models, but their widespread adoption has been limited by the computational burden of training them. We propose a novel loss function called Energy Discrepancy (ED) which does not rely on the computation of scores or expensive Markov chain Monte Carlo. We show that ED approaches the explicit score matching and negative log-likelihood loss under different limits, effectively interpolating between both. Consequently, minimum ED estimation overcomes the problem of nearsightedness encountered in score-based estimation methods, while also enjoying theoretical guarantees. Through numerical experiments, we demonstrate that ED learns low-dimensional data distributions faster and more accurately than explicit score matching or contrastive divergence. For high-dimensional image data, we describe how the manifold hypothesis puts limitations on our approach and demonstrate the effectiveness of energy discrepancy by training the energy-based model as a prior of a variational decoder model.
Code (1)
Tasks
DecoderSimilar Papers 제목 키워드 기반
Graph Size-imbalanced Learning with Energy-guided Structural Smoothing
Graph is a prevalent data structure employed to represent the relationships between entities, frequently serving as a tool to depict and simulate numerous systems, such as molecules and social networks. However, real-wor…
Graph ClassificationTrajectory-Independent Flexibility Envelopes of Energy-Constrained Systems with State-Dependent Losses
As non-dispatchable renewable power units become prominent in electric power grids, demand-side flexibility appears as a key element of future power systems' operation. Power and energy bounds are intuitive metrics to de…
Mixability of Integral Losses: a Key to Efficient Online Aggregation of Functional and Probabilistic Forecasts
In this paper we extend the setting of the online prediction with expert advice to function-valued forecasts. At each step of the online game several experts predict a function, and the learner has to efficiently aggrega…
Measuring Heterogeneity in Machine Learning with Distributed Energy Distance
In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensi…
Federated LearningExploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
The diverse nature of dialects presents challenges for models trained on specific linguistic patterns, rendering them susceptible to errors when confronted with unseen or out-of-distribution (OOD) data. This study introd…
Dialect IdentificationOut-of-Distribution Detection