Wasserstein PAC-Bayes Learning: Exploiting Optimisation Guarantees to Explain Generalisation
PAC-Bayes learning is an established framework to both assess the generalisation ability of learning algorithms, and design new learning algorithm by exploiting generalisation bounds as training objectives. Most of the exisiting bounds involve a \emph{Kullback-Leibler} (KL) divergence, which fails to capture the geometric properties of the loss function which are often useful in optimisation. We address this by extending the emerging \emph{Wasserstein PAC-Bayes} theory. We develop new PAC-Bayes bounds with Wasserstein distances replacing the usual KL, and demonstrate that sound optimisation guarantees translate to good generalisation abilities. In particular we provide generalisation bounds for the \emph{Bures-Wasserstein SGD} by exploiting its optimisation properties.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Rigorous Link between Deep Ensembles and (Variational) Bayesian Methods
We establish the first mathematically rigorous link between Bayesian, variational Bayesian, and ensemble methods. A key step towards this it to reformulate the non-convex optimisation problem typically encountered in dee…
Deep LearningUncertainty QuantificationVariational InferenceEnsemble Distributionally Robust Bayesian Optimisation with Continuous Context
We study Bayesian Optimisation (BO) in settings where the objective function is influenced by uncontrollable environmental contexts governed by an unknown probability distribution. In practice, the contextual distributio…
Bayesian Optimisation with Formal Guarantees
Application domains of Bayesian optimization include optimizing black-box functions or very complex functions. The functions we are interested in describe complex real-world systems applied in industrial settings. Even t…
Bayesian OptimisationBayesian OptimizationBures-Wasserstein Importance-Weighted Evidence Lower Bound: Exposition and Applications
The Importance-Weighted Evidence Lower Bound (IW-ELBO) has emerged as an effective objective for variational inference (VI), tightening the standard ELBO and mitigating the mode-seeking behaviour. However, optimizing the…
Tree-Wasserstein Barycenter for Large-Scale Multilevel Clustering and Scalable Bayes
We study in this paper a variant of Wasserstein barycenter problem, which we refer to as tree-Wasserstein barycenter, by leveraging a specific class of ground metrics, namely tree metrics, for Wasserstein distance. Drawi…
Clustering