paper-with-me

Papers

Target Layer Regularization for Continual Learning Using Cramer-Wold Generator

2021-11-15 · Marcin Mazur, Łukasz Pustelnik, Szymon Knop, Patryk Pagacz, Przemysław Spurek

We propose an effective regularization strategy (CW-TaLaR) for solving continual learning problems. It uses a penalizing term expressed by the Cramer-Wold distance between two probability distributions defined on a target layer of an underlying neural network that is shared by all tasks, and the simple architecture of the Cramer-Wold generator for modeling output data representation. Our strategy preserves target layer distribution while learning a new task but does not require remembering previous tasks' datasets. We perform experiments involving several common supervised frameworks, which prove the competitiveness of the CW-TaLaR method in comparison to a few existing state-of-the-art continual learning models.

📄 PDF Abstract BibTeX arXiv:2111.07928

Code (1)

gmum/cw-talar 공식 구현 pytorch

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Non-linear ICA based on Cramer-Wold metric

2019-03-01 · Przemysław Spurek, Aleksandra Nowak, Jacek Tabor, Łukasz Maziarka 외

Non-linear source separation is a challenging open problem with many applications. We extend a recently proposed Adversarial Non-linear ICA (ANICA) model, and introduce Cramer-Wold ICA (CW-ICA). In contrast to ANICA we u…

Balanced Marginal and Joint Distributional Learning via Mixture Cramer-Wold Distance

2023-12-06 · SeungHwan An, Sungchul Hong, Jong-June Jeon

In the process of training a generative model, it becomes essential to measure the discrepancy between two high-dimensional probability distributions: the generative distribution and the ground-truth distribution of the …

Synthetic Data Generation

Cramer-Wold AutoEncoder

2018-05-23 · ICLR 2019 5 · Szymon Knop, Jacek Tabor, Przemysław Spurek, Igor Podolak 외

We propose a new generative model, Cramer-Wold Autoencoder (CWAE). Following WAE, we directly encourage normality of the latent space. Our paper uses also the recent idea from Sliced WAE (SWAE) model, which uses one-dime…

Joint Distributional Learning via Cramer-Wold Distance

2023-10-25 · SeungHwan An, Jong-June Jeon

The assumption of conditional independence among observed variables, primarily used in the Variational Autoencoder (VAE) decoder modeling, has limitations when dealing with high-dimensional datasets or complex correlatio…

DecoderSynthetic Data Generation

SeGMA: Semi-Supervised Gaussian Mixture Auto-Encoder

2019-06-21 · Marek Śmieja, Maciej Wołczyk, Jacek Tabor, Bernhard C. Geiger

We propose a semi-supervised generative model, SeGMA, which learns a joint probability distribution of data and their classes and which is implemented in a typical Wasserstein auto-encoder framework. We choose a mixture …

Style Transfer