Multifactor Sequential Disentanglement via Structured Koopman Autoencoders
Disentangling complex data to its latent factors of variation is a fundamental task in representation learning. Existing work on sequential disentanglement mostly provides two factor representations, i.e., it separates the data to time-varying and time-invariant factors. In contrast, we consider multifactor disentanglement in which multiple (more than two) semantic disentangled components are generated. Key to our approach is a strong inductive bias where we assume that the underlying dynamics can be represented linearly in the latent space. Under this assumption, it becomes natural to exploit the recently introduced Koopman autoencoder models. However, disentangled representations are not guaranteed in Koopman approaches, and thus we propose a novel spectral loss term which leads to structured Koopman matrices and disentanglement. Overall, we propose a simple and easy to code new deep model that is fully unsupervised and it supports multifactor disentanglement. We showcase new disentangling abilities such as swapping of individual static factors between characters, and an incremental swap of disentangled factors from the source to the target. Moreover, we evaluate our method extensively on two factor standard benchmark tasks where we significantly improve over competing unsupervised approaches, and we perform competitively in comparison to weakly- and self-supervised state-of-the-art approaches. The code is available at https://github.com/azencot-group/SKD.
Code (1)
Tasks
DisentanglementInductive BiasRepresentation LearningSimilar Papers 제목 키워드 기반
On Disentanglement in Gaussian Process Variational Autoencoders
Complex multivariate time series arise in many fields, ranging from computer vision to robotics or medicine. Often we are interested in the independent underlying factors that give rise to the high-dimensional data we ar…
DisentanglementTime SeriesTime Series AnalysisKoopman Regularized Deep Speech Disentanglement for Speaker Verification
Human speech contains both linguistic content and speaker dependent characteristics making speaker verification a key technology in identity critical applications. Modern deep learning speaker verification systems aim to…
Representation LearningSpeaker VerificationContext-Enhanced CSI Tracking Using Koopman-Inspired Dual Autoencoders in Dynamic Wireless Environments
This paper introduces a novel framework for tracking and predicting Channel State Information (CSI) by leveraging Physics-Informed Autoencoders (PIAE) integrated with a learned Koopman operator. The proposed approach mod…
Computational EfficiencyPrivacy PreservingLearning the Koopman Operator using Attention Free Transformers
Learning Koopman operators with autoencoders enables linear prediction in a latent space, but long-horizon rollouts often drift off the learned manifold, leading to phase and amplitude errors on systems with switching, c…
Forecasting Sequential Data using Consistent Koopman Autoencoders
Recurrent neural networks are widely used on time series data, yet such models often ignore the underlying physical structures in such sequences. A new class of physics-based methods related to Koopman theory has been in…
Time SeriesTime Series Analysis